<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dishant Sharma</title>
    <description>The latest articles on DEV Community by Dishant Sharma (@dishant0406).</description>
    <link>https://dev.to/dishant0406</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F811279%2Fed56a090-b880-4fb2-b8a2-d988877fc75f.png</url>
      <title>DEV Community: Dishant Sharma</title>
      <link>https://dev.to/dishant0406</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dishant0406"/>
    <language>en</language>
    <item>
      <title>Jev by TypeSafe AI: the hype, reactions, and two-week clone war</title>
      <dc:creator>Dishant Sharma</dc:creator>
      <pubDate>Sun, 04 Oct 2026 11:52:23 +0000</pubDate>
      <link>https://dev.to/dishant0406/jev-by-typesafe-ai-the-hype-reactions-and-two-week-clone-war-132f</link>
      <guid>https://dev.to/dishant0406/jev-by-typesafe-ai-the-hype-reactions-and-two-week-clone-war-132f</guid>
      <description>&lt;p&gt;Somewhere in a roleplay community, people are feeding their own chat replies into a model built for industrial automation. The model scores their writing and tells them which reply sounds more in character. That model is Jev, and it is not what TypeSafe AI put in its press release.&lt;/p&gt;

&lt;p&gt;i found this at 1 in the morning, scrolling thread titles i did not expect. Jev came out of stealth on September 15, 2026. TypeSafe AI is founded by Diogo Almeida, who helped build the instruction-following work inside OpenAI that became ChatGPT.&lt;/p&gt;

&lt;p&gt;The pitch is simple. Jev never writes text. You send program state, ask typed questions, and get back probabilities in 70 to 500 milliseconds. Output tokens cost nothing.&lt;/p&gt;

&lt;p&gt;The reaction split in a familiar way. r/LocalLLaMA ran 330 comments calling it old tech in a new costume. Someone in r/BetterOffline put it more simply:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Jev is Jian Yang's Hotdog / Not Hotdog app.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is a Silicon Valley joke, and it landed.&lt;/p&gt;

&lt;p&gt;And then the odd part. The busiest threads were in r/SillyTavernAI, a roleplay tool, and r/SideProject, people scoring their startup ideas. Two weeks later, OpenAI shipped a Decisions API built on Luna, and AWS shipped a local model called Strands Decider 2B.&lt;/p&gt;

&lt;p&gt;TechCrunch ran the headline "Amazon releases its own Jev clone as decision models flood the web."&lt;/p&gt;

&lt;p&gt;Here's why you should care. If the next phase of AI is models that stop talking, the way we build agents changes. And i don't think anyone has fully processed that yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Jev actually is
&lt;/h2&gt;

&lt;p&gt;Here's a question people keep asking: is Jev another chatbot? No. Jev is a discriminative model. It classifies and scores. It does not generate.&lt;/p&gt;

&lt;p&gt;You give it a state, usually text or JSON, plus typed questions. It answers all of them in one call, in parallel. The whole vocabulary fits in three primitives:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;noul&lt;/code&gt;: a yes or no question, answered as a 0 to 1 probability.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;choice&lt;/code&gt;: pick from a fixed list, each option gets a probability.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;score&lt;/code&gt;: rate the input against a numerical scale or rubric.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is the entire API surface.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;There is no free text anywhere.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;i kept seeing Jev mentioned and assumed it was a chat model. A thread on r/PromptEngineering said the same, and the top reply accused the poster of advertising. The confusion is fair. We are not used to a model that outputs nothing but verdicts.&lt;/p&gt;

&lt;p&gt;The endpoint is a single line: &lt;code&gt;api.typesafe.ai/v1/systemone&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why 200 milliseconds matters
&lt;/h2&gt;

&lt;p&gt;What actually happens is that Jev skips token generation. Normal models emit one token at a time, each conditioned on the last. Jev samples every answer at once, in parallel, hardware aware. That is where the speed comes from.&lt;/p&gt;

&lt;p&gt;The training is different too. RLHF optimizes for what human raters prefer. RLCD, short for Reinforcement Learning for Calibrated Decisions, optimizes for honest probabilities. High confidence means high accuracy, and the model is built so you can trust the number.&lt;/p&gt;

&lt;p&gt;That matters more than speed. If a model can do a task 95 percent of the time but cannot say when it is in the failing 5 percent, you cannot wire it into anything.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Jev&lt;/th&gt;
&lt;th&gt;Chat LLM&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;typed decisions&lt;/td&gt;
&lt;td&gt;free text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency&lt;/td&gt;
&lt;td&gt;70 to 500 ms&lt;/td&gt;
&lt;td&gt;3 to 329 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Confidence&lt;/td&gt;
&lt;td&gt;calibrated&lt;/td&gt;
&lt;td&gt;often overconfident&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The skeptics have a point
&lt;/h2&gt;

&lt;p&gt;The loudest take on r/LocalLLaMA is that Jev is not new. A classifier with calibrated probabilities is a solved problem. TypeSafe has not published enough architecture detail to verify the interesting claims.&lt;/p&gt;

&lt;p&gt;i could be wrong here, but the architecture silence bothers me more than the price. The company itself admits it cannot prove the pricing is not subsidized.&lt;/p&gt;

&lt;p&gt;Input tokens cost $0.042 per million. Output tokens are free. That price is a bet, not a fact.&lt;/p&gt;

&lt;p&gt;Jev also has a documented gap. It does not have deep knowledge of niche domains. You supply the context, which means you are doing some of the work. &lt;strong&gt;That is fine for routing. It is not fine for expert judgment.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Nobody expected the response
&lt;/h2&gt;

&lt;p&gt;i used to think a model needed years to matter. Jev got cloned in two weeks.&lt;/p&gt;

&lt;p&gt;OpenAI announced its Decisions API ahead of DevDay, built on the small Luna model, about 150 milliseconds per call. AWS shipped Strands Decider 2B, small enough to run locally. Two giants, two weeks, two takes on the same idea.&lt;/p&gt;

&lt;p&gt;The strangest adoption is the roleplay one. People wired Jev into SillyTavern to score replies. Average response around 200 milliseconds, at roughly $0.0005 per call. The most human use of a machine-native model i have seen.&lt;/p&gt;

&lt;p&gt;And builders went further. There is Claude Code memory tooling built on Jev. Idea scoring side projects with 130 comments. Cloudflare Workers running it. Everyone is bolting a verdict machine onto whatever they already made.&lt;/p&gt;




&lt;h2&gt;
  
  
  A quick detour about naming
&lt;/h2&gt;

&lt;p&gt;All of this traces back to one book. Jev is named after Daniel Kahneman's System 1, the fast intuitive thinking from Thinking, Fast and Slow.&lt;/p&gt;

&lt;p&gt;Every AI cycle, someone rediscovers that book and names a product after it. We survived years of System 2 reasoning marketing, and now the pendulum swings to System 1.&lt;/p&gt;

&lt;p&gt;i do the same thing. There is a script in my home directory called "yesterday" because i was reading Murakami when i wrote it.&lt;/p&gt;

&lt;p&gt;It has nothing to do with time travel. It parses CSV files.&lt;/p&gt;

&lt;p&gt;Naming is a mood ring.&lt;/p&gt;

&lt;p&gt;The funny part: Kahneman thought fast thinking was the source of bias. The confident automatic judgments were the ones to distrust. Half of the Jev hype is celebrating exactly that, which means we are misusing the book again.&lt;/p&gt;

&lt;p&gt;But that argument can wait.&lt;/p&gt;




&lt;h2&gt;
  
  
  The honest version
&lt;/h2&gt;

&lt;p&gt;Most people do not need Jev. If you make under a few hundred decisions a day, a chat model with a careful prompt works fine.&lt;/p&gt;

&lt;p&gt;A normal classifier works fine too, as long as your labels do not change often. The calibrated probabilities only matter at volume, where a silent wrong answer costs real money.&lt;/p&gt;

&lt;p&gt;And this is a bet. The pricing looks subsidized, and the company says so. The architecture is underpublished.&lt;/p&gt;

&lt;p&gt;The clones are already here, and a decision model is easier to copy than a frontier model. If OpenAI's and AWS's versions are good, the moat is thin.&lt;/p&gt;

&lt;p&gt;"cannot hallucinate" is also narrower than it sounds. Jev cannot type nonsense because it cannot type. &lt;strong&gt;It can still decide wrong, confidently, and your pipeline will trust it.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Last thought
&lt;/h2&gt;

&lt;p&gt;i keep thinking about the roleplay people. A model built to route transactions and guardrail agents.&lt;/p&gt;

&lt;p&gt;The most boring enterprise tool imaginable, and the community that adopted it fastest was people grading each other's fiction.&lt;/p&gt;

&lt;p&gt;That is the part nobody predicted. The enterprise will take years to trust calibrated probabilities. The weird corners of the internet plugged it in and shipped.&lt;/p&gt;

&lt;p&gt;Usually it works the other way around.&lt;/p&gt;

&lt;p&gt;So here is the question i keep circling. When your model stops talking, what changes about how you build?&lt;/p&gt;

&lt;p&gt;And the sharper one: what were you using the chat for in the first place?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>news</category>
      <category>startup</category>
    </item>
    <item>
      <title>Kimi K2.7-Code is Out. The Open-Source Coding Model That Thinks Less</title>
      <dc:creator>Dishant Sharma</dc:creator>
      <pubDate>Sun, 04 Oct 2026 09:14:08 +0000</pubDate>
      <link>https://dev.to/dishant0406/kimi-k27-code-is-out-the-open-source-coding-model-that-thinks-less-pi3</link>
      <guid>https://dev.to/dishant0406/kimi-k27-code-is-out-the-open-source-coding-model-that-thinks-less-pi3</guid>
      <description>&lt;p&gt;A post in r/singularity with the title "Kimi 2.7 code is released &amp;amp; open-sourced" hit 100 upvotes fast. Over on X, the announcement from @Kimi_Moonshot crossed 2,600 likes while i was refreshing the page. This is the third major open-source model release from Moonshot AI in six months. K2.5 landed in January. K2.6 dropped in April. And now we have K2.7-Code.&lt;/p&gt;

&lt;p&gt;That pace is the part that hits me. Not the numbers yet. The rhythm.&lt;/p&gt;

&lt;p&gt;Most AI labs ship a model, take a victory lap, and you hear from them again six to nine months later. Moonshot is running on a different clock. Every two months, something new lands on Hugging Face with open weights and a permissive license. And each time, the gap to the proprietary frontier models gets a little thinner.&lt;/p&gt;

&lt;p&gt;So what did they actually ship this time?&lt;/p&gt;

&lt;h2&gt;
  
  
  The model in numbers
&lt;/h2&gt;

&lt;p&gt;Kimi-K2.7-Code is a coding-focused agentic model built on top of K2.6. Same architecture underneath. 1 trillion total parameters, 32 billion activated, 256K context window, Mixture-of-Experts with 384 experts and MLA attention. The vision encoder stayed too. But the benchmarks tell a different story from the architecture sheet.&lt;/p&gt;

&lt;p&gt;Here's how it stacks up against its predecessor and the competition:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;K2.6&lt;/th&gt;
&lt;th&gt;K2.7 Code&lt;/th&gt;
&lt;th&gt;GPT-5.5&lt;/th&gt;
&lt;th&gt;Claude Opus 4.8&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Kimi Code Bench v2&lt;/td&gt;
&lt;td&gt;50.9&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;62.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;69.0&lt;/td&gt;
&lt;td&gt;67.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Program Bench&lt;/td&gt;
&lt;td&gt;48.3&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;53.6&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;69.1&lt;/td&gt;
&lt;td&gt;63.8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MLS Bench Lite&lt;/td&gt;
&lt;td&gt;26.7&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;35.1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;35.5&lt;/td&gt;
&lt;td&gt;42.8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi Claw 24/7 Bench&lt;/td&gt;
&lt;td&gt;42.9&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;46.9&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;52.8&lt;/td&gt;
&lt;td&gt;50.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MCP Atlas&lt;/td&gt;
&lt;td&gt;69.4&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;76.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;79.4&lt;/td&gt;
&lt;td&gt;81.3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MCP Mark Verified&lt;/td&gt;
&lt;td&gt;72.8&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;81.1&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;92.9&lt;/td&gt;
&lt;td&gt;76.4&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The biggest jump is MLS Bench Lite. That's a 31.5% improvement over K2.6. MLS Bench tests whether AI systems can invent generalizable ML methods. It's hard. K2.7 went from looking okay to sitting right next to GPT-5.5 on that one.&lt;/p&gt;

&lt;p&gt;But here's what caught my attention more than the gains.&lt;/p&gt;

&lt;h2&gt;
  
  
  Less overthinking
&lt;/h2&gt;

&lt;p&gt;Every reasoning model right now is on a trajectory toward longer and longer thinking chains. More tokens. More deliberation. More time spent in the model's head before it types anything out. GPT-5.5 does it. Claude Opus 4.8 does it. DeepSeek's models do it. The implicit assumption is that more thinking equals better results.&lt;/p&gt;

&lt;p&gt;Kimi K2.7 goes the other way.&lt;/p&gt;

&lt;p&gt;Moonshot claims 30% lower reasoning-token usage compared to K2.6. They call it "less overthinking." That phrase is doing a lot of work. It suggests that the previous models were thinking more than they needed to, and that trimming that fat did not hurt performance. The numbers back up the claim. Most benchmarks went up despite fewer thinking tokens.&lt;/p&gt;

&lt;p&gt;Think about what this means for a real coding session. You ask K2.7 to refactor a module. Instead of spending 3,000 reasoning tokens debating whether to use a factory pattern, it picks one and moves on. The code comes out right. You pay for fewer tokens. Everyone wins.&lt;/p&gt;

&lt;p&gt;i think about the engineering behind that. A model that gets to the right answer faster is not just cheaper to run. It is better at handling long conversations and complex multi-step tasks. The reasoning chain does not grow forever. It stays focused. That matters more than people give it credit for.&lt;/p&gt;

&lt;h2&gt;
  
  
  The long-horizon angle
&lt;/h2&gt;

&lt;p&gt;The other headline improvement is long-horizon coding. End-to-end task success rates went up. That means K2.7 is better at taking a complex request, working through it step by step, and delivering the finished product without getting lost halfway.&lt;/p&gt;

&lt;p&gt;This is where coding models usually fall apart. They nail the first function but forget the overall architecture. They write great tests but miss the integration point. K2.7's improvements here feel like the real gain, even if the in-house benchmarks are harder to verify independently.&lt;/p&gt;

&lt;p&gt;Another detail worth mentioning. The model forces "preserve thinking" mode. That means the reasoning content is kept across multi-turn interactions. For coding agents that need context from earlier in the conversation, this is a practical feature that makes a real difference in how the model behaves.&lt;/p&gt;

&lt;p&gt;The weights and code are on Hugging Face right now under a Modified MIT License. You can pull them, run them on vLLM or SGLang, and start building. The API is available at platform.moonshot.ai for people who do not want to self-host.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why i think about naming schemes
&lt;/h2&gt;

&lt;p&gt;This has nothing to do with benchmarks. But i have to say it.&lt;/p&gt;

&lt;p&gt;Moonshot AI names their models like Apple names iOS versions. K2.1, K2.5, K2.6, K2.7. Each one is a point release. That feels weird for a 1-trillion-parameter model. You expect major version jumps for something this big. But the more i think about it, the more it makes sense. These are not separate research projects. They are iterative improvements on the same architecture. The point numbers reflect that honestly.&lt;/p&gt;

&lt;p&gt;Most AI companies would have called K2.7 "Kimi-4" or "Kimi-Ultra" or something with a trademark symbol. Moonshot just called it K2.7. There is something refreshing about that lack of marketing nonsense. Just a model, a number, and a link to the weights.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it falls short
&lt;/h2&gt;

&lt;p&gt;Kimi K2.7 is not better than GPT-5.5 or Claude Opus 4.8 on most benchmarks. That is the honest truth. Look at the table. GPT-5.5 leads on Kimi Code Bench v2, Program Bench, MCP Atlas, MCP Mark Verified, and Kimi Claw 24/7. Claude wins on MLS Bench Lite. K2.7 does not hold the top spot on a single metric.&lt;/p&gt;

&lt;p&gt;But raw benchmark scores are not the whole story. K2.7 is open-source. You can run it on your own hardware. You can fine-tune it. You can audit the weights. The cost per token is dramatically lower than GPT-5.5 or Claude. For teams building production systems that need predictable costs and data privacy, those considerations often outweigh a 5-10% benchmark gap.&lt;/p&gt;

&lt;p&gt;The honest take: if you are building a coding agent that needs the absolute best performance and your budget allows API pricing from Anthropic or OpenAI, go with Claude or GPT. If you want something open, customizable, and good enough to handle most real coding tasks, K2.7 is the strongest option from an open-source lab right now.&lt;/p&gt;

&lt;h2&gt;
  
  
  One more thing
&lt;/h2&gt;

&lt;p&gt;The 6x High-Speed Mode they mentioned in the announcement is coming soon. Not available yet. That feels like the kind of feature that could shift the calculus further when it lands. A six times speed boost on top of 30% less thinking tokens would make this model genuinely fast. Fast enough that latency-sensitive applications become viable.&lt;/p&gt;

&lt;p&gt;But it is not here yet. So we wait.&lt;/p&gt;

&lt;p&gt;i keep coming back to the release cadence. K2.5 in January. K2.6 in April. K2.7 in June. Moonshot is not slowing down. At this rate, K2.8 could be here before the end of summer. And the gap to the frontier will be even smaller. Or gone.&lt;/p&gt;

&lt;p&gt;There is a lesson in here somewhere. Maybe it is that open-source AI moves faster when you stop trying to make every release a revolution. Maybe it is that the labs that keep shipping eventually catch up. Or maybe it is just that the second half of 2026 is going to be very interesting for anyone who cares about coding models.&lt;/p&gt;

</description>
      <category>kimi</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Inside Claude Fable 5: The Beast, the Limiter, the Fallout</title>
      <dc:creator>Dishant Sharma</dc:creator>
      <pubDate>Sun, 04 Oct 2026 09:14:06 +0000</pubDate>
      <link>https://dev.to/dishant0406/inside-claude-fable-5-the-beast-the-limiter-the-fallout-4e0e</link>
      <guid>https://dev.to/dishant0406/inside-claude-fable-5-the-beast-the-limiter-the-fallout-4e0e</guid>
      <description>&lt;p&gt;Simon Willison spent 5.5 hours throwing everything he had at it.&lt;/p&gt;

&lt;p&gt;His verdict: "it's a beast."&lt;/p&gt;

&lt;p&gt;Across the same hours, another phrase was circulating in the same threads: "a Ferrari with a 30mph limiter."&lt;/p&gt;

&lt;p&gt;Both are describing the same model. That tension is the whole story of Claude Fable 5.&lt;/p&gt;

&lt;p&gt;Anthropic dropped Fable 5 on June 9, 2026. It's the first Mythos-class model they have let the public touch.&lt;/p&gt;

&lt;p&gt;The numbers are absurd. SWE-bench Pro at 80.3, GPT-5.5 at 58.6. On Cognition's FrontierCode Diamond, Fable scores 29.3% against Opus 4.8 at 13.4% and GPT-5.5 at 5.7%.&lt;/p&gt;

&lt;p&gt;That is not a small gap. That is a different tier.&lt;/p&gt;

&lt;p&gt;i have been watching the reactions roll in for two days. The developer community is split in a way i have not seen since GPT-4 launched. Not between fans and critics. Between people who tried it and people who read about it.&lt;/p&gt;

&lt;p&gt;Both are reacting to something real.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the numbers actually mean
&lt;/h2&gt;

&lt;p&gt;Stripe reported that Fable 5 migrated a 50-million-line Ruby codebase in a day. Their estimate for a team doing it by hand was over two months.&lt;/p&gt;

&lt;p&gt;Cursor called it their best result on CursorBench. They said it "opened up a class of long-horizon problems that were out of reach." Replit said the same thing in fewer words. Less time, fewer tokens.&lt;/p&gt;

&lt;p&gt;Here is what that looks like in practice. On "high" thinking mode, Fable produces better results than Opus 4.8 on "xhigh." Large refactors that used to hit context limits just finish now.&lt;/p&gt;

&lt;p&gt;Bugs that Opus missed get caught. And it does this while using fewer tokens per task.&lt;/p&gt;

&lt;p&gt;But fewer tokens per task does not mean fewer dollars. The pricing is $10 per million input tokens and $50 per million output tokens. That is double Opus 4.8. Complex sessions regularly run 500k to 1 million tokens. At $50 per million output, a single serious session can cost real money.&lt;/p&gt;

&lt;p&gt;i saw someone on HN break it down cleanly. If Fable lands the answer in one pass where Opus needs four, the math is the same. Several teams say their daily TCO went down because they stopped paying for retries. Others say it went up because they changed nothing and doubled their per-call cost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The truth is both groups are right. It depends entirely on your use case.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A debate broke out on HN the same day. Someone said benchmarks do not matter, vibes are all that count. Someone else fired back that you cannot run a lab on vibes. Both were serious. Both made good points.&lt;/p&gt;

&lt;p&gt;Here is where it gets messy. The day after launch, Fortune ran a story with the headline "Anthropic accused of 'secret sabotage.'" The phrase is dramatic.&lt;/p&gt;

&lt;p&gt;The content is worse.&lt;/p&gt;

&lt;p&gt;Fable 5 has a feature buried in its 319-page system card. When the model detects a request related to AI development work, it silently downgrades the quality of its response. The user does not get a notification. The model just gets worse at the thing you asked it to do.&lt;/p&gt;

&lt;p&gt;And it goes back to normal for the next query.&lt;/p&gt;

&lt;p&gt;Anthropic says this affects about 0.03% of traffic. But the system card says something specific: this restriction is "not visible to the user." The model still responds. It just uses "interventions to limit Claude's effectiveness" without telling you.&lt;/p&gt;

&lt;p&gt;The reaction was immediate. Nathan Lambert called it "appalling" and said it paints Anthropic as "anti-science." Dean Ball said it "massively and profoundly raises the status of the argument that AI safety has been hype to justify monopolistic behavior." Jeremy Howard said Anthropic is "allowing themselves to use their top model for frontier AI research" while sabotaging others who try.&lt;/p&gt;

&lt;p&gt;Behnam Neyshabur, who used to co-lead Anthropic's AI scientist effort, posted: "Working on AI for cancer? Sorry, i can't help you. Working on AI for Alzheimer's? Sorry, i'm becoming a bit dumb when it comes to the AI part of it."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This is the part that stuck with me. Not the policy debate. The fact that a former Anthropic scientist is saying this publicly.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The safeguards story
&lt;/h2&gt;

&lt;p&gt;The cybersecurity classifiers are genuinely impressive. Fable 5 complied with zero harmful single-turn requests in testing. External red teams found it was the toughest model they tested.&lt;/p&gt;

&lt;p&gt;Over 1,000 hours of bug bounty work produced no universal jailbreaks. The UK AISI made some progress, but only within a brief initial window.&lt;/p&gt;

&lt;p&gt;Biology and chemistry requests fall back to Opus 4.8. The model can predict viral shell assembly properties better than dedicated protein language models. That is dual-use capability in plain view.&lt;/p&gt;

&lt;p&gt;You cannot have that power without some controls. But here is the distinction that matters. When Fable blocks a cybersecurity query, it tells you. It falls back to Opus 4.8 and you see the message. When Fable blocks an AI research query, it does not tell you. It just gets subtly worse at the task, and you have to figure it out yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  I think about beer sometimes
&lt;/h2&gt;

&lt;p&gt;There is a brewery near my apartment that makes an IPA with a cult following. The head brewer once told me they intentionally make the first batch of each season weaker than the recipe calls for. Their reasoning: they want people to try it, like it, come back, and then get the real version later. If they sold the full strength version first, people would complain it was too much.&lt;/p&gt;

&lt;p&gt;i thought about that story while reading the Fable system card. Not because the situations are the same. They are not. But because both involve a maker deciding the consumer cannot handle the full version yet. And both involve not telling them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who should actually care about this
&lt;/h2&gt;

&lt;p&gt;If you are building a startup with Claude Code and your agentic sessions cost $20 each instead of $10 but finish in half the prompts, you probably should not care about the sabotage debate. The model is better. Your costs might even go down. The controversy is about 0.03% of traffic.&lt;/p&gt;

&lt;p&gt;If you are doing AI research, you should care a lot. The model that is best at your work is now deliberately worse at your work when it detects you doing it. And it does not tell you. That is a real problem for scientific progress.&lt;/p&gt;

&lt;p&gt;If you are just using Claude for writing, analysis, or coding side projects, this launch is almost certainly good news. Fable is better than Opus at nearly everything. The price increase matters, but the capability increase matters more.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The honest take: most people do not need Fable 5. Opus 4.8 is still excellent and costs half as much.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  One more thing
&lt;/h2&gt;

&lt;p&gt;i keep coming back to the same question. Not whether Fable 5 is good. It is obviously good. The question is whether a model this capable and this restricted can stay this way. The "Ferrari with a 30mph limiter" comparison is clever but incomplete. A real Ferrari with a limiter is still fun to drive. You feel the power underneath. You know it is there.&lt;/p&gt;

&lt;p&gt;With Fable, you do not know when the limiter kicks in. You just get a worse result and move on. And that is the part that makes this launch different from every other model release this year. Not the benchmark scores. Not the pricing. The fact that the most capable model in the world is also the least transparent about what it is actually doing.&lt;/p&gt;

</description>
      <category>claude</category>
    </item>
    <item>
      <title>Grok Build Hype vs Reality: A Look at Real User Reactions</title>
      <dc:creator>Dishant Sharma</dc:creator>
      <pubDate>Sun, 04 Oct 2026 09:13:32 +0000</pubDate>
      <link>https://dev.to/dishant0406/grok-build-hype-vs-reality-a-look-at-real-user-reactions-48hg</link>
      <guid>https://dev.to/dishant0406/grok-build-hype-vs-reality-a-look-at-real-user-reactions-48hg</guid>
      <description>&lt;p&gt;"Grok Build feels like the previous generation of coding models."&lt;/p&gt;

&lt;p&gt;That's from a Reddit post on r/grok. From someone who actually paid $99 to try SuperGrok Heavy.&lt;/p&gt;

&lt;p&gt;Another commenter chimed in with their own story. Three hours trying Grok Build on an existing codebase, watching it silently change behavior without a warning. Not a wrong output.&lt;/p&gt;

&lt;p&gt;A silent behavioral shift that the user only caught by accident.&lt;/p&gt;

&lt;p&gt;And then the kicker: the terminal won't even let you copy-paste error messages.&lt;/p&gt;

&lt;p&gt;This is the reception for xAI's big coding agent play. Grok Build landed on May 25, 2026 as an early beta. A terminal-native CLI that directly challenges Claude Code and Codex CLI.&lt;/p&gt;

&lt;p&gt;The hype was loud. Elon posted. The xAI blog went up.&lt;/p&gt;

&lt;p&gt;A bunch of tech outlets ran the headline. But the real story is what happened after people actually installed it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;So what is everyone arguing about?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Grok Build is not a chat interface or a VS Code plugin. It runs in your terminal. You point it at a project, describe what you want, and it plans, searches your codebase, writes code, and shows diffs for review.&lt;/p&gt;

&lt;p&gt;The install is one line: &lt;code&gt;curl -fsSL https://x.ai/cli/install.sh | bash&lt;/code&gt;. You authenticate with your xAI account and you're in.&lt;/p&gt;

&lt;p&gt;The default mode is Plan Mode. Before touching a single file, Grok Build proposes a step-by-step plan. You approve it, comment on individual steps, or rewrite it entirely.&lt;/p&gt;

&lt;p&gt;Nothing runs until you sign off. This directly fixes the thing developers hate most about coding agents.&lt;/p&gt;

&lt;p&gt;Here's the problem it solves: an agent does something wrong, and by the time you notice, three other things have already changed downstream. Plan Mode puts a checkpoint between "task given" and "codebase modified."&lt;/p&gt;

&lt;p&gt;Claude Code does not have this natively. That is a genuine edge.&lt;/p&gt;

&lt;p&gt;But the architecture goes deeper. Grok Build spawns up to eight parallel subagents, each working in its own Git worktree. They do not step on each other.&lt;/p&gt;

&lt;p&gt;An evaluation layer called Arena Mode scores competing outputs before you review. Larger refactors that would take one agent an hour can be parallelized across several agents working in isolation.&lt;/p&gt;

&lt;p&gt;It also ships with MCP support. It picks up AGENTS.md conventions. And it supports headless mode via the &lt;code&gt;-p&lt;/code&gt; flag for CI pipelines.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The gap nobody is glossing over&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;xAI published a score of 70.8% on SWE-Bench Verified for the model powering Grok Build. Claude Opus 4.7 Adaptive sits at 87.6%. OpenAI's Codex is at 85%.&lt;/p&gt;

&lt;p&gt;That is a 17-point gap. Not a rounding error. Not a methodology disagreement.&lt;/p&gt;

&lt;p&gt;xAI's response is that benchmarks don't fully reflect real-world engineering. That is technically true of every benchmark. It does not close a 17-point gap.&lt;/p&gt;

&lt;p&gt;On simple, scoped tasks the difference may be invisible. On complex multi-file work, it shows up as more failed attempts, more reverted diffs, and more human review time.&lt;/p&gt;

&lt;p&gt;One reviewer reported hallucinated edits under heavy load. Corrupted Dockerfiles from ambiguous prompts.&lt;/p&gt;

&lt;p&gt;The competitive context makes it worse. Codex has passed three million weekly active users. Claude Code has driven Anthropic to $30 billion in annual recurring revenue.&lt;/p&gt;

&lt;p&gt;Grok Build enters with none of that production history.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And then there's the pricing&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The pricing structure is where things get interesting. Access started at $300 per month for SuperGrok Heavy. Then they opened it to SuperGrok at $30 and X Premium Plus at $40, with a promotional tier at $99.&lt;/p&gt;

&lt;p&gt;Claude Code costs $20 flat.&lt;/p&gt;

&lt;p&gt;On HN, the sentiment was blunt. "Only $300 a month. (Or $3,000 a year.) The xAI casino wants all your money even if you don't use it for a month."&lt;/p&gt;

&lt;p&gt;Another user: "I'm not spending $300 a month on something my employer will never approve me to use."&lt;/p&gt;

&lt;p&gt;The API pricing is more reasonable. $1 per million input tokens and $2 per million output tokens. And a Grok Build 0.1 model is available through OpenRouter and Vercel AI Gateway.&lt;/p&gt;

&lt;p&gt;But the subscription model for the CLI itself is polarizing. At $30, it's competitive. At $300, it's a tough sell against Claude Code at $20.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The thing that actually surprised me&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The TUI is genuinely impressive. I expected something slapped together. But one of xAI's engineers confirmed on HN that the terminal interface is written in Rust using Ratatui.&lt;/p&gt;

&lt;p&gt;Proper vim keybindings. Mouse support. Careful alt-screen rendering.&lt;/p&gt;

&lt;p&gt;They put real work into making the terminal experience feel polished.&lt;/p&gt;

&lt;p&gt;And the binary is local-first. Your source code, credentials, and project data stay on your machine. They don't get transmitted to xAI's servers for every operation.&lt;/p&gt;

&lt;p&gt;For a company that could easily justify cloud-only processing, that's a meaningful choice.&lt;/p&gt;

&lt;p&gt;But here's the detail that made me pause: someone pointed out you can't copy-paste errors from the Grok Build terminal. A basic UX gap in a tool that wants to replace your existing coding workflow.&lt;/p&gt;

&lt;p&gt;Small things like this are why "v0.1" matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I actually think about coding agents in 2026&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;i spent an afternoon digging through HN threads, Reddit posts, and review sites to understand the reception. And what i found is that the conversation around Grok Build is less about Grok Build itself and more about where we are with coding agents in general.&lt;/p&gt;

&lt;p&gt;Everyone is tired of benchmarks. Everyone is tired of announcements. What people actually want is a tool that doesn't silently break their codebase.&lt;/p&gt;

&lt;p&gt;And Grok Build, for all its architectural ambition, is a v0.1 that needs the model to catch up to the interface.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The architecture is ahead of the model. That's the honest take.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you're on SuperGrok already, try it. It costs nothing extra. But if you're deciding between Claude Code at $20 and Grok Build at $30, the benchmark gap is real and you will feel it on complex tasks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Speaking of beautiful terminal tools&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This whole research detour reminded me of the time i spent a weekend building a TUI for a side project using Bubble Tea in Go. Spent two days getting the layout right, another day on keybindings, and on Monday realized nobody would ever use it.&lt;/p&gt;

&lt;p&gt;i built it for a problem nobody had.&lt;/p&gt;

&lt;p&gt;That's kind of where Grok Build is right now. Beautiful terminal, real architectural thinking, and a model that needs maybe one more training run to justify the hype.&lt;/p&gt;

&lt;p&gt;But at least Ratatui produces gorgeous terminal UIs. i could say the same about my Bubble Tea experiment. Except nobody used it.&lt;/p&gt;

&lt;p&gt;We'll see if people actually use Grok Build.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who should actually care about this&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Let me be direct.&lt;/p&gt;

&lt;p&gt;Individual dev paying out of pocket? Grok Build at $30 is worth trying. At $300, it is not.&lt;/p&gt;

&lt;p&gt;The model quality is not there yet. You will get more value from Claude Code at $20 for the same tasks.&lt;/p&gt;

&lt;p&gt;Team on xAI infrastructure or with SuperGrok enterprise access? The parallel subagent architecture is genuinely interesting. Plan Mode is a real differentiator.&lt;/p&gt;

&lt;p&gt;But you need to accept that the underlying model will produce more bad outputs than Claude or Codex. Your team will spend more time reviewing.&lt;/p&gt;

&lt;p&gt;Building agent orchestration tools? The ACP support matters. The headless mode and API access mean you can build custom workflows on top of Grok Build in a way that is harder with Claude Code's more limited automation surface.&lt;/p&gt;

&lt;p&gt;For everyone else: wait for the next model iteration. The architecture is promising.&lt;/p&gt;

&lt;p&gt;The model needs to catch up. That is the honest assessment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One last thing&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;i keep thinking about that Reddit comment. Three hours watching a tool silently change code behavior. The user caught it by accident.&lt;/p&gt;

&lt;p&gt;Not because the tool warned them. But because they happened to look at the diff before committing.&lt;/p&gt;

&lt;p&gt;That is the real problem Grok Build needs to solve. Not the SWE-Bench score. Not the pricing.&lt;/p&gt;

&lt;p&gt;The trust gap.&lt;/p&gt;

&lt;p&gt;A terminal agent that quietly makes things worse while looking like it's helping is worse than no agent at all.&lt;/p&gt;

&lt;p&gt;Plan Mode is a step in the right direction. Parallel agents are cool.&lt;/p&gt;

&lt;p&gt;But trust is earned one clean diff at a time.&lt;/p&gt;

&lt;p&gt;And right now, Grok Build is still earning it.&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
    <item>
      <title>open-slide Hit 4k Stars. The Reactions Told a Real Story</title>
      <dc:creator>Dishant Sharma</dc:creator>
      <pubDate>Sun, 04 Oct 2026 09:13:30 +0000</pubDate>
      <link>https://dev.to/dishant0406/open-slide-hit-4k-stars-the-reactions-told-a-real-story-3i3l</link>
      <guid>https://dev.to/dishant0406/open-slide-hit-4k-stars-the-reactions-told-a-real-story-3i3l</guid>
      <description>&lt;p&gt;Yiwei Ho posted every slide at the Rayboba event was built with open-slide. Not a demo. Not a mockup.&lt;/p&gt;

&lt;p&gt;Decks for a community event hosted by Raycast. And the slides looked clean.&lt;/p&gt;

&lt;p&gt;This matters because slide tools built for agents have become their own category this year. There is AI presentation generators everywhere. Type your topic, get a PDF.&lt;/p&gt;

&lt;p&gt;They all promise the same thing. And they all deliver the same disappointment.&lt;/p&gt;

&lt;p&gt;Bullet points on a gradient background. A logo in the corner. Something that looks like a presentation but feels like a template you have seen a thousand times.&lt;/p&gt;

&lt;p&gt;open-slide does not do that. And people noticed.&lt;/p&gt;

&lt;p&gt;The GitHub repo hit around 4,000 stars in about a week. The creator posted a video of Cursor building a deck in under a minute.&lt;/p&gt;

&lt;p&gt;212 likes, 16,000 views. Replies poured in.&lt;/p&gt;

&lt;p&gt;"So cool!" "This is awesome." "Yo lfg."&lt;/p&gt;

&lt;p&gt;But also real questions. "Is it PPTX compatible?" "Can I export to PowerPoint?"&lt;/p&gt;

&lt;p&gt;The excitement was loud. So was the hesitation.&lt;/p&gt;

&lt;p&gt;I spent last weekend reading every reply. Wanted to understand what made this different from the other AI slide tools I have tried. And I found something I did not expect.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It is not a slide generator. It is a slide framework.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What open-slide actually is
&lt;/h2&gt;

&lt;p&gt;open-slide is a React-first framework. Every slide is a React component on a 1920x1080 canvas. No templates. No drag and drop.&lt;/p&gt;

&lt;p&gt;No limited layouts.&lt;/p&gt;

&lt;p&gt;You describe your deck in natural language. Your coding agent writes the React. open-slide handles the canvas, navigation, hot reload, and present mode.&lt;/p&gt;

&lt;p&gt;You start with one command.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @open-slide/cli init my-deck
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That scaffolds a workspace. Then you use &lt;code&gt;/create-slide&lt;/code&gt; with your agent.&lt;/p&gt;

&lt;p&gt;The skill asks four questions. Topic and aesthetic. Page count.&lt;/p&gt;

&lt;p&gt;Text density. Motion or static.&lt;/p&gt;

&lt;p&gt;Based on your answers, it plans the structure and writes the pages as React components.&lt;/p&gt;

&lt;p&gt;The first time I read this, I thought it sounded like extra steps. Why not just use a prompt and get a PDF?&lt;/p&gt;

&lt;p&gt;But here is the difference. A PDF is dead. A React component is alive.&lt;/p&gt;

&lt;p&gt;You can edit it, inspect it, change anything.&lt;/p&gt;

&lt;p&gt;I have tried the AI presentation generators. NotebookLM gave me the closest thing to a real deck.&lt;/p&gt;

&lt;p&gt;But every edit meant rerunning the whole prompt and hoping for the best. That is not editing. That is gambling.&lt;/p&gt;

&lt;p&gt;open-slide works differently. The output is not a black box. It is a file in your project.&lt;/p&gt;

&lt;p&gt;You can open it in your editor. Fix a typo. Swap a color.&lt;/p&gt;

&lt;p&gt;The agent owns the creation. You own the file.&lt;/p&gt;

&lt;h2&gt;
  
  
  The inspect and comment loop
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;This is the part that got me.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In the dev server, you can click any element and attach a comment. "Make this red." "Change the headline." "Shrink this font."&lt;/p&gt;

&lt;p&gt;Those comments get saved as markers in the source code. Then you run &lt;code&gt;/apply-comments&lt;/code&gt; and the agent edits the files. It clears the markers when done.&lt;/p&gt;

&lt;p&gt;The loop is simple. Present. Click to comment. Apply.&lt;/p&gt;

&lt;p&gt;Repeat.&lt;/p&gt;

&lt;p&gt;No back and forth about what "section 3" means. No regenerating the whole deck for one color change. You point at the thing. The agent changes the thing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happened on X
&lt;/h2&gt;

&lt;p&gt;The reactions split into two groups.&lt;/p&gt;

&lt;p&gt;One group was excited. Vinay Juneja called it "damn fast and soo crzy shipping man." Alexi said "Someone has been on fire lately!"&lt;/p&gt;

&lt;p&gt;The other group asked practical questions.&lt;/p&gt;

&lt;p&gt;fil asked "is it pptx compatible?" Elias Lumer asked "Can u export these slides into an edit-able PowerPoint?"&lt;/p&gt;

&lt;p&gt;These are not complaints. They are constraints. Most people do not deliver from a terminal. They need a file to share. A format their team can open.&lt;/p&gt;

&lt;p&gt;But one reply stood out. vinaykumar wrote: "Decks are a good stress test because layout exposes bad agent state fast. Harder to fake than a clean text diff."&lt;/p&gt;

&lt;p&gt;This is the honest take. open-slide tests how well your agent handles layout. And layout is where most agents fall apart.&lt;/p&gt;

&lt;p&gt;There is another angle here. open-slide works with any agent. Claude Code. Codex. Cursor. Gemini CLI.&lt;/p&gt;

&lt;p&gt;The framework does not care which one you use. Each agent gets the same /create-slide skill and the same canvas rules. The quality difference between agents shows up in the deck.&lt;/p&gt;

&lt;p&gt;That is the stress test.&lt;/p&gt;

&lt;p&gt;One command exports the whole deck as a static HTML site or a print-ready PDF. No server needed. One click deploy to Vercel, Cloudflare Pages, Netlify.&lt;/p&gt;

&lt;p&gt;This is the part I think gets overlooked. You are not locked into anything. The output is plain static files.&lt;/p&gt;

&lt;p&gt;Here is what it does:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Agent-native authoring with built-in skills&lt;/li&gt;
&lt;li&gt;In-browser inspector with comment editing&lt;/li&gt;
&lt;li&gt;Assets manager with svgl logo search&lt;/li&gt;
&lt;li&gt;Present mode with speaker notes and timer&lt;/li&gt;
&lt;li&gt;Export to static HTML and PDF&lt;/li&gt;
&lt;li&gt;Slide manager with folders and drag and drop&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here is what it does not do yet:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Export to PPTX&lt;/li&gt;
&lt;li&gt;Drag and drop editing for non developers&lt;/li&gt;
&lt;li&gt;Work without an agent&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The thing I keep thinking about
&lt;/h2&gt;

&lt;p&gt;I spent an embarrassing amount of time searching for "open-slide" and getting results for OpenSlide. You know, the medical image viewer. Thousands of pathology slide scans.&lt;/p&gt;

&lt;p&gt;It has been around for years.&lt;/p&gt;

&lt;p&gt;This happens every time I name a project. You search for something obvious and find three other things with the same name.&lt;/p&gt;

&lt;p&gt;I once named a side project "Tracer" and found twelve SaaS products using the same name. Maddening.&lt;/p&gt;

&lt;p&gt;open-slide is stuck with this for now. The SEO overlap is real.&lt;/p&gt;

&lt;p&gt;But maybe that changes if the project keeps growing. The star history graph is vertical. That kind of growth helps.&lt;/p&gt;

&lt;p&gt;I also wonder about the name itself. "open-slide" sounds like it describes what it is. Open slide framework.&lt;/p&gt;

&lt;p&gt;But it also sounds like open-source slide. And that is accurate. It is MIT licensed.&lt;/p&gt;

&lt;p&gt;You can fork it. Modify it. Build your own workflows on top of it.&lt;/p&gt;

&lt;p&gt;That matters more to me than the name confusion.&lt;/p&gt;

&lt;p&gt;But I still wish naming things was easier.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest talk about who should use this
&lt;/h2&gt;

&lt;p&gt;Most people do not need this.&lt;/p&gt;

&lt;p&gt;If your presentations have eight slides and bullet points, use Google Slides. It is faster. Your team can edit it.&lt;/p&gt;

&lt;p&gt;No one needs React for a status update.&lt;/p&gt;

&lt;p&gt;If you work in a corporate environment where everything must be a .pptx file, this tool is not for you yet. The export question came up multiple times in replies.&lt;/p&gt;

&lt;p&gt;That is a gap.&lt;/p&gt;

&lt;p&gt;open-slide makes sense when your deck has visual density. Charts. Custom layouts. Animations.&lt;/p&gt;

&lt;p&gt;When you want each slide to feel intentional instead of templated. And when you already have an agent setup ready.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do not pick up a React framework for slides unless you have the agent workflow figured out.&lt;/strong&gt; The tool is not the hard part. The agent workflow is.&lt;/p&gt;

&lt;h2&gt;
  
  
  A final thought
&lt;/h2&gt;

&lt;p&gt;I keep coming back to that one reply. "Decks are a good stress test because layout exposes bad agent state fast."&lt;/p&gt;

&lt;p&gt;This is the signal. Not the 4,000 stars. Not the landing page.&lt;/p&gt;

&lt;p&gt;open-slide reveals where agent coding is right now. It can handle the hard part. Layout.&lt;/p&gt;

&lt;p&gt;Visual design. Complex positioning.&lt;/p&gt;

&lt;p&gt;The part that usually gives agents away as robots.&lt;/p&gt;

&lt;p&gt;But it also reveals the gap. People want to share these files. People want to collaborate.&lt;/p&gt;

&lt;p&gt;People want a format that works outside the dev environment.&lt;/p&gt;

&lt;p&gt;I think about the Rayboba event. Real slides. Real audience.&lt;/p&gt;

&lt;p&gt;Real presentation. That is more than most "AI slide tools" have achieved.&lt;/p&gt;

&lt;p&gt;What is the point of a perfect slide if you cannot email it to your boss?&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>agents</category>
    </item>
    <item>
      <title>Two Prompts Was All It Took to Ditch Claude Design for Open Source</title>
      <dc:creator>Dishant Sharma</dc:creator>
      <pubDate>Sun, 04 Oct 2026 09:12:57 +0000</pubDate>
      <link>https://dev.to/dishant0406/two-prompts-was-all-it-took-to-ditch-claude-design-for-open-source-46hm</link>
      <guid>https://dev.to/dishant0406/two-prompts-was-all-it-took-to-ditch-claude-design-for-open-source-46hm</guid>
      <description>&lt;p&gt;Someone on Reddit signed up for a Claude subscription to try Claude Design. They hit their weekly quota after two prompts. Two prompts.&lt;/p&gt;

&lt;p&gt;The thread was full of people nodding along. The hype was enormous. The limits were real.&lt;/p&gt;

&lt;p&gt;Claude Design launched to massive noise. The demos looked incredible. Full web prototypes from a single prompt.&lt;/p&gt;

&lt;p&gt;Mobile flows. Slide decks with WebGL backgrounds. But people hit a wall fast.&lt;/p&gt;

&lt;p&gt;Cloud-only. Locked to Anthropic's model. A paid plan that runs out of quota way too fast.&lt;/p&gt;

&lt;p&gt;One user on r/ClaudeAI described it as a tool crippled by bugs with a brutal usage limit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That is when Open Design showed up.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The nexu-io team released Open Design as the open source, local-first alternative. Apache-2.0 license. Your own agent and your own API key.&lt;/p&gt;

&lt;p&gt;Within days it had 18,000 stars. Now it sits at over 55,900 on GitHub. Better Stack made a video asking "Why 40k Developers Abandoned Claude Design."&lt;/p&gt;

&lt;p&gt;The timing was not a coincidence.&lt;/p&gt;

&lt;h3&gt;
  
  
  What it actually does
&lt;/h3&gt;

&lt;p&gt;Open Design turns your coding agent into a design engine. It auto-detects the CLI agents you already have. Claude Code, Cursor, Codex, Gemini, Qwen, OpenCode.&lt;/p&gt;

&lt;p&gt;12 adapters now, and it detects them automatically on boot.&lt;/p&gt;

&lt;p&gt;Clone the repo, run pnpm install, then pnpm tools-dev run web. A local SQLite daemon and web UI spin up. No cloud and no subscription.&lt;/p&gt;

&lt;p&gt;You can run it on a VPS or your laptop.&lt;/p&gt;

&lt;p&gt;Here is what ships:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;139 skill bundles (file-based SKILL.md files)&lt;/li&gt;
&lt;li&gt;150 portable DESIGN.md systems for Linear, Stripe, Vercel, Apple&lt;/li&gt;
&lt;li&gt;MCP integration so your editor reads design files directly&lt;/li&gt;
&lt;li&gt;BYOK at every layer, any OpenAI-compatible API&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The MCP integration is what people keep coming back to. A Reddit user called it "the part worth paying attention to." Drop it into Cursor or Windsurf and your editor reads your design files.&lt;/p&gt;

&lt;p&gt;No more copy-pasting code or taking screenshots for your agent.&lt;/p&gt;

&lt;h3&gt;
  
  
  The debate that would not die
&lt;/h3&gt;

&lt;p&gt;The most interesting part of the Reddit threads was not the tool. It was the argument that followed. Someone called the post out for calling Open Design "free" when you still need a good model to get results.&lt;/p&gt;

&lt;p&gt;Then it exploded.&lt;/p&gt;

&lt;p&gt;"It's free as in speech, not free as in beer," one comment said.&lt;/p&gt;

&lt;p&gt;"No, the tool is free. The inference is not," said another.&lt;/p&gt;

&lt;p&gt;Dozens of comments went back and forth. People arguing semantics while the project sat there with 55,000 stars and working code. One user pointed out that even a 32B local model was not enough.&lt;/p&gt;

&lt;p&gt;Another replied "you can say this about every self hosted LLM alternative."&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The real divide is between people who think "free" means no cost at all and people who understand Open Design is a wrapper for models you already pay for.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The comment that stuck with me was someone comparing it to a free drill bit. "Is a free drill bit not free because you need a drill?" That is where this conversation landed.&lt;/p&gt;

&lt;p&gt;But there is a real point buried in the noise. A lot of people want local-first AI tools that work with local models. And right now local models are just not good enough for design work.&lt;/p&gt;

&lt;p&gt;That is not Open Design's fault. It is a hardware problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  But the quality gap is real
&lt;/h3&gt;

&lt;p&gt;The honest feedback from people who tried it is mixed. flickerdown said it straight: "It's not great. Cannot generate even remotely the same quality level as Claude Designer."&lt;/p&gt;

&lt;p&gt;biglboy tried it with Kimi K2.6, DeepSeek V4 Pro, and Gemini 3.1 Flash. "Everything is looking really, really, really bad. It builds it but it is just awful."&lt;/p&gt;

&lt;p&gt;The project is at v0.8.0. Rough edges are expected. Surgical edits are on the roadmap.&lt;/p&gt;

&lt;p&gt;Some users hit errors loading preview frames. Getting Ollama or llama.cpp to connect has also been hard for some.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;But Claude Design has its own problems.&lt;/strong&gt; Users describe it as having "more bad days than good," being "fucking awful for exporting," and breaking into a "black screen of death."&lt;/p&gt;

&lt;p&gt;The gap between them is not as wide as people think.&lt;/p&gt;

&lt;p&gt;One user pointed out that Open Design lets you use models like Kimi K2.6. Those are "10x cheaper and very close to Opus 4.7" for design work.&lt;/p&gt;

&lt;p&gt;The value equation depends heavily on which model you pair with the tool.&lt;/p&gt;

&lt;h3&gt;
  
  
  What the architecture buys you
&lt;/h3&gt;

&lt;p&gt;The BYOK approach is what makes this interesting long term. Models shift every few weeks. A new model drops, a better one appears, prices change.&lt;/p&gt;

&lt;p&gt;Locked into Claude Design? You wait for Anthropic. Use Open Design and plug in the new model.&lt;/p&gt;

&lt;p&gt;A Reddit comment put it well: "BYOK is the part that matters long term. Models shift every six weeks and getting locked into one vendor's prompt format is a real footgun."&lt;/p&gt;

&lt;p&gt;The 139 skills are also portable. Drop a folder into the skills directory and restart the daemon. A new capability appears with no plugin store or approval process.&lt;/p&gt;

&lt;p&gt;No vendor gatekeeping.&lt;/p&gt;

&lt;p&gt;Prototype with Gemini Flash or a local Ollama setup. Switch to Claude Opus or Kimi K2.6 for final polish. That flexibility matters when you build daily.&lt;/p&gt;

&lt;p&gt;Another detail that stood out: you can export projects from Claude Design as a ZIP and drag them into Open Design. The import path exists. That is smart.&lt;/p&gt;




&lt;h3&gt;
  
  
  Speaking of switching models
&lt;/h3&gt;

&lt;p&gt;I spent an embarrassing amount of time this week trying to connect my local setup. Typed the base URL wrong. The error said "invalid or unreachable" and I stared at it for 20 minutes.&lt;/p&gt;

&lt;p&gt;Then I realized the port was off by one.&lt;/p&gt;

&lt;p&gt;This is the kind of nonsense that happens with local-first tools. It is frustrating. But it is also mine.&lt;/p&gt;

&lt;p&gt;No one can throttle my port number. No one can decide I hit my limit and cut me off. I just need to type the right port.&lt;/p&gt;




&lt;h3&gt;
  
  
  The real talk
&lt;/h3&gt;

&lt;p&gt;If you expect Open Design to match Claude Design's output quality today, you will be disappointed. It does not. Not even close, based on user reports.&lt;/p&gt;

&lt;p&gt;The v0.8 release is ambitious but the polish is not there.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Most people should stick with Claude Design if they can afford the quota.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;But here is the thing. Claude Design's value falls apart fast if you build more than a couple things a day. The quota hits, the black screens appear, and the export fails.&lt;/p&gt;

&lt;p&gt;And you cannot fix any of it because you do not own the tool.&lt;/p&gt;

&lt;p&gt;Open Design wins on ownership. Swap models, edit the skills, run it offline, and contribute back. That matters for building real products, not just trying demos.&lt;/p&gt;

&lt;p&gt;The architecture is right even if the current output is not.&lt;/p&gt;

&lt;p&gt;I keep thinking about that Reddit user who hit quota after two prompts. They paid for a subscription. They wanted to try the thing everyone hyped.&lt;/p&gt;

&lt;p&gt;They got two uses before the door shut. Two is not a number that inspires confidence.&lt;/p&gt;

&lt;p&gt;Open Design is not there yet. But it is local, it is open, and it is yours. When the next model drops and Claude Design's output still looks better, you just plug it in and keep going.&lt;/p&gt;

&lt;p&gt;That is the whole point.&lt;/p&gt;

</description>
      <category>claude</category>
      <category>design</category>
    </item>
    <item>
      <title>The oh-my-pi Hype: Hashline, Reactions, and What Everyone Missed</title>
      <dc:creator>Dishant Sharma</dc:creator>
      <pubDate>Sun, 04 Oct 2026 09:12:55 +0000</pubDate>
      <link>https://dev.to/dishant0406/the-oh-my-pi-hype-hashline-reactions-and-what-everyone-missed-2771</link>
      <guid>https://dev.to/dishant0406/the-oh-my-pi-hype-hashline-reactions-and-what-everyone-missed-2771</guid>
      <description>&lt;p&gt;There's a number floating around that i cannot stop thinking about. Grok 4 Fast went from 6.7% to 68.3% on coding edits. No new model release. No training data dump.&lt;/p&gt;

&lt;p&gt;Just a better tool format. The creator of oh-my-pi built something called Hashline. Every line of code gets a short content hash, and the model references the hash instead of reproducing the entire line. Tenfold improvement in an afternoon.&lt;/p&gt;

&lt;p&gt;This is the hashline effect. And it tells you everything about why the coding agent space is so weird right now.&lt;/p&gt;

&lt;p&gt;oh-my-pi is a fork of Mario Zechner's Pi project, extended by a developer called can1357. It is a terminal-based AI coding agent that runs on Bun, with a Rust core clocking around 27k lines.&lt;/p&gt;

&lt;p&gt;It supports 40+ model providers, 32 built-in tools, 13 LSP operations, and 27 DAP operations. A lot of people are talking about it. And the reactions are all over the place.&lt;/p&gt;

&lt;p&gt;So what is actually happening here.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The hype is real but nobody agrees on why&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Some people see oh-my-pi as the LazyVim of coding agents. Batteries included, works out of the box. You don't have to spend your weekend configuring tools.&lt;/p&gt;

&lt;p&gt;Others call it bloated and say the real way is pure Pi, building your own setup, composing your own distro.&lt;/p&gt;

&lt;p&gt;AkitaOnRails tested it and said: "It works. It is not bad. It also did not make me say how did I live without this." That is a pretty honest take from someone who has used a lot of these tools.&lt;/p&gt;

&lt;p&gt;On Reddit, a user called it "really nice in that it feels like a polished product, but the initial token count is 22k, which hurts." And that is a real problem. Every conversation starts with a 22k token overhead before you even say hello to the model.&lt;/p&gt;

&lt;p&gt;But then you have can1357's numbers. Grok Code Fast jumping from 6.7% to 68.3%. Gemini improving by 8 percentage points. Grok 4 Fast spending 61% fewer output tokens.&lt;/p&gt;

&lt;p&gt;These are not small wins.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What makes oh-my-pi different&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The tool surface is the actual product here. Most coding agents give the model a Python sandbox and call it done. oh-my-pi runs a persistent Python kernel and a Bun worker, both of which can call back into the agent's own tools over a loopback bridge. So the agent can load a CSV with the read tool from inside Python, chart it from JavaScript, and never leave the cell.&lt;/p&gt;

&lt;p&gt;The LSP integration is serious too. Ask for a rename and it goes through workspace/willRenameFiles. Re-exports, barrel files, and aliased imports all update before the file moves. That is the kind of thing you only get if someone actually shipped a product, not just a script.&lt;/p&gt;

&lt;p&gt;And the DAP support is wild. A C binary segfaults. The agent attaches lldb, steps to the bad pointer, reads the frame. A Go service hangs. DlV, walk the goroutines.&lt;/p&gt;

&lt;p&gt;A Python process is wedged. Debugpy, pause, inspect, evaluate. Most agents are still sprinkling print statements.&lt;/p&gt;

&lt;p&gt;i used to think the model was everything. Pick the biggest one, throw tokens at it, get results. oh-my-pi is making me rethink that.&lt;/p&gt;

&lt;p&gt;The difference between a 6.7% pass rate and a 68.3% pass rate is not a bigger brain. It is a better interface.&lt;/p&gt;

&lt;p&gt;Here is a question people always ask: is oh-my-pi better than Claude Code? And the answer is it depends on what you value. Claude Code has the best context management and compaction in the game. The subscription pricing is unbeatable if you use it all day.&lt;/p&gt;

&lt;p&gt;oh-my-pi gives you flexibility. 40+ providers. Custom tools. Debugger integration. Things Claude Code does not let you touch.&lt;/p&gt;

&lt;p&gt;But honesty requires saying this: oh-my-pi feels like a lab project in some places. The default verbosity is annoying. The configuration surface is deep and not always documented well.&lt;/p&gt;

&lt;p&gt;Akita put it best when he said he did not want to spend an afternoon tuning prompts and panel sizes.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;the model is the moat. the harness is the bridge. burning bridges just means fewer people bother to cross.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That quote from can1357 is the key insight. Models are getting better every quarter. The gap between them is shrinking. What actually determines whether an agent helps you or not is the harness around it.&lt;/p&gt;

&lt;p&gt;Here's what broke for one HN user who migrated from Pi to oh-my-pi:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the Rust tools using NAPI caused tool confusion after updates&lt;/li&gt;
&lt;li&gt;models needed 11 iterations just to edit a single file&lt;/li&gt;
&lt;li&gt;the agent then gave up and rewrote the whole file from memory&lt;/li&gt;
&lt;li&gt;fixing it meant wiping all agent memory and starting fresh&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The integrations with search, LSP, syntax highlighting, and custom tools work like a charm, they said. But the upgrade path is not smooth.&lt;/p&gt;

&lt;p&gt;And this is the tension. oh-my-pi adds so much. But every layer you add is another thing that can break when you pull from upstream.&lt;/p&gt;

&lt;p&gt;Here is something related but off topic. The naming situation is hilarious. Pi is already a famously overloaded name. Raspberry Pi, the math constant, and now a coding agent.&lt;/p&gt;

&lt;p&gt;And oh-my-pi adds another layer. Type "oh my pi" into Google and you get a mixture of coding agent docs and confused home assistant users. Every time someone talks about their "Pi setup" in a coding context, they have to clarify which Pi.&lt;/p&gt;

&lt;p&gt;i have been in conversations where three different people assumed three different things. It is the namespace collision of 2026 and nobody has fixed it yet.&lt;/p&gt;

&lt;p&gt;But really this reminds me of NeoVim. Some people want LazyVim. Some people want to build from scratch. Neither is wrong.&lt;/p&gt;

&lt;p&gt;The problem is when people on either side pretend their choice is the only right one.&lt;/p&gt;

&lt;p&gt;And the irony is that this debate over batteries included versus minimal keeps happening with every new tool. It happened with text editors, with frameworks, with package managers.&lt;/p&gt;

&lt;p&gt;Every time the community splits into two camps that spend more energy fighting each other than building. oh-my-pi is just the latest frontier for this. &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who should actually use oh-my-pi&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here is the blunt version. If you use Claude Code all day and it works fine for you, oh-my-pi probably is not worth switching to. The subscription pricing from Anthropic changes the math completely. Why pay per token when you can pay per month.&lt;/p&gt;

&lt;p&gt;But if you want to use models outside the big two. If you want to run local models. If you want the full tool surface including LSP, DAP, persistent kernels, subagent orchestration, and hash-anchored edits.&lt;/p&gt;

&lt;p&gt;If you care about the harness as much as the model. Then oh-my-pi is probably the most complete open option right now.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The hashline thing still bothers me&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;i keep going back to that number. 6.7% to 68.3%. Not a better model. Not more training compute.&lt;/p&gt;

&lt;p&gt;Just a better way to say "edit this line." The old approach was: the model guesses the exact text around the change, repeats it, and hopes for the best. Line numbers shift, whitespace mismatches happen.&lt;/p&gt;

&lt;p&gt;The edit fails, the model retries, the context grows, your token bill doubles.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Reddit thread said it best&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;On r/PiCodingAgent, someone asked "Is oh-my-pi the best batteries-included Pi mod, or does it diverge from Pi's minimalist philosophy too much?" The top comment: "The main goal of pi is not to be minimalistic but to be customized by the end-user with no built-in trash."&lt;/p&gt;

&lt;p&gt;That is a very specific philosophy. And oh-my-pi violates it on purpose. It adds everything. It is the opposite of minimalist.&lt;/p&gt;

&lt;p&gt;Hashline: every line has a hash. The model says "edit this hash." The tool finds it. Done.&lt;/p&gt;

&lt;p&gt;This is the kind of improvement that changes how you think about the whole space. Maybe the next 10x is not a bigger model. Maybe it is a smarter way to talk to the one you already have.&lt;/p&gt;

&lt;p&gt;i still think about that HN comment comparing Pi to llama-cpp. "The former gets all the hype, while the latter is actually the more impressive part."&lt;/p&gt;

&lt;p&gt;It is true for oh-my-pi too. The hype cycle is pointing at the model. The real work is happening in the harness. And the numbers say it matters more than most people are willing to admit.&lt;/p&gt;

</description>
      <category>coding</category>
      <category>tooling</category>
    </item>
    <item>
      <title>Step 3.7 Flash: The 198B MoE Model Everyone Is Actually Running</title>
      <dc:creator>Dishant Sharma</dc:creator>
      <pubDate>Sun, 04 Oct 2026 09:12:22 +0000</pubDate>
      <link>https://dev.to/dishant0406/step-37-flash-the-198b-moe-model-everyone-is-actually-running-pb1</link>
      <guid>https://dev.to/dishant0406/step-37-flash-the-198b-moe-model-everyone-is-actually-running-pb1</guid>
      <description>&lt;p&gt;someone on X posted a photo of a DGX Spark sitting on a regular desk with a terminal window running step 3.7 flash. a 198 billion parameter vision model. on a box that fits next to a monitor. and it was not a flex post. it was a "here is the config that saved me three hours" post.&lt;/p&gt;

&lt;p&gt;that is the kind of energy around this release.&lt;/p&gt;

&lt;p&gt;stepfun dropped step 3.7 flash on may 29. two days ago as i write this. and the AI community has been processing it ever since. not because it is the biggest model ever made. but because of how much it does for how little it costs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;a 198B model that acts like an 11B model at inference time.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;here is what that means in practice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The architecture that makes the cost work&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;step 3.7 flash is a sparse mixture of experts model. 198 billion total parameters but only about 11 billion activate per token. the other 187 billion just sit there waiting for the right input. this is not new. MoE is well understood by now. but the execution matters.&lt;/p&gt;

&lt;p&gt;the model activates 8 out of 288 experts per token. the routing decides which experts fire based on what the input actually needs. so the compute cost per token looks much closer to a dense 11B model than a 198B one. the result is up to 400 tokens per second throughput. that is genuinely fast for a model with this capability profile.&lt;/p&gt;

&lt;p&gt;on the vision side, stepfun added a 1.8B ViT encoder on top of the language backbone. that means native image understanding without the full multimodal tax that some larger models pay. it handles UI wireframes, charts, dense documents, and natural scenes. then it writes code or calls tools based on what it sees.&lt;/p&gt;

&lt;p&gt;the vision results back this up. SimpleVQA with search tools hits 79.2, first place in that category, ahead of GPT 5.5 at 79.1. V* with Python tools reaches 95.3, competitive with Kimi K2.6 at 96.9 and Gemini 3 Flash at 96.3. these are flash tier results matching pro tier models on visual tasks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why people are excited&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;the X feed has been full of people actually running this thing. not theorizing. not reviewing benchmarks from a blog post. loading it on hardware and seeing what happens.&lt;/p&gt;

&lt;p&gt;sudo su posted his exact llama.cpp config for running it on a DGX Spark. someone else posted benchmark comparisons with claude opus 4.6. multiple people noted that step 3.5 flash was already their most used model for agentic coding. now 3.7 is out and it is a meaningful jump.&lt;/p&gt;

&lt;p&gt;the numbers that keep getting shared:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ClawEval-1.1: 67.1 (leads the benchmark, next best is 59.8)&lt;/li&gt;
&lt;li&gt;SWE-Bench Verified with Advisor Mode: 76.3% at $0.19 per task&lt;/li&gt;
&lt;li&gt;Claude Opus 4.6: 78.7% at $1.76 per task&lt;/li&gt;
&lt;li&gt;SimpleVQA with Search: 79.2 (beat GPT 5.5 at 79.1)&lt;/li&gt;
&lt;li&gt;BrowseComp: 75.82% (close to Claude Opus 4.7 at 79.3%)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;that is 97% of claude opus 4.6's coding performance at one ninth the cost.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;people are not excited because it beats everything. they are excited because it changes the math on what you can put in production.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Advisor Mode is the actual trick&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;here is what makes the cost number possible.&lt;/p&gt;

&lt;p&gt;the basic problem with long agentic runs is that most of the work is routine. tool calls, reading results, iterating on straightforward steps. you do not need a frontier model for that. but occasionally the agent hits a genuinely hard decision point - a planning step that could send the whole trajectory wrong, or a recovery from repeated failures.&lt;/p&gt;

&lt;p&gt;advisor mode keeps step 3.7 flash in control of the full run. it calls tools, reads results, and iterates end to end. but at specific inflection points where its own judgment is not sufficient, it consults a larger advisor model and then continues executing on that guidance.&lt;/p&gt;

&lt;p&gt;the expensive model gets used only where it actually matters. the small model stays cheap through most of the run. the advisor cost gets amortized across many steps rather than paid on every token.&lt;/p&gt;

&lt;p&gt;stepfun describes this as their implementation of the executor-advisor strategy that anthropic wrote about. it is not a new idea. but the execution and the pricing make it work in practice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where it actually falls short&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;terminal-bench 2.1 is the clearest gap. step 3.7 flash scores 59.5 against GPT 5.5 at 82.7 and gemini 3.5 flash at 76.2. that is not a close race. for workflows that depend heavily on terminal interaction and complex command execution, this model is not the right choice right now.&lt;/p&gt;

&lt;p&gt;GDPval across 44 occupations shows 45.8% against GPT 5.5 and claude opus 4.7 at 63%. that is a significant gap for general professional task coverage.&lt;/p&gt;

&lt;p&gt;HLE with tools at 47.2% trails Claude Opus 4.7 at 54.7% and GPT 5.5 at 52.2%. for the hardest reasoning tasks with tool use, frontier models still have a clear edge.&lt;/p&gt;

&lt;p&gt;and all benchmark numbers here are from stepfun's own evaluation unless otherwise noted. self-reported results should be read with that caveat.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That time I spent 2 hours reading the wrong docs&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;speaking of caveats, i once spent 2 hours reading the wrong version of an API doc because the url had "v2" but the page said "v3". the entire time i was confused about why the endpoints did not match. turns out v2 was the current version and v3 was the upcoming one. the docs were published early by accident.&lt;/p&gt;

&lt;p&gt;this happens to me more than i want to admit. which is why when i see a model release with clear deployment instructions and multiple quantization formats on day one, i pay attention. stepfun shipped BF16, FP8, NVFP4, and GGUF weights on huggingface simultaneously. that is not nothing. that is a team that knows people will actually try to run this thing.&lt;/p&gt;

&lt;p&gt;and they did. people are running it on DGX Spark, on mac studios with 128GB unified memory, on cloud instances through nim and openrouter. the ecosystem support at launch matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The honest take&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;most people do not need step 3.7 flash.&lt;/p&gt;

&lt;p&gt;if you are building a chatbot that answers questions from a knowledge base, a smaller dense model will serve you better. if you need the absolute best terminal agent performance, frontier models are still ahead. if you have fewer than 10k requests a day, the pricing difference does not matter enough to justify the infrastructure complexity.&lt;/p&gt;

&lt;p&gt;but if you are building production agentic workflows where tool orchestration reliability and cost per task matter more than raw benchmark ceiling, this model is worth serious evaluation. the ClawEval lead and the advisor mode cost profile are both genuinely differentiated.&lt;/p&gt;

&lt;p&gt;the cross-harness consistency is also underrated. step 3.7 flash narrowed per-harness variance from 43-73% (step 3.5) down to 64.5-71.5%. that means more predictable agent behavior across different scaffolds. in production, that matters more than a 2% benchmark gain.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I keep thinking about&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;i keep coming back to that X post. someone running a 198B vision model on a desktop machine. sharing their exact config so others do not have to spend 3 hours debugging.&lt;/p&gt;

&lt;p&gt;that is the real story here. not the benchmark scores. not the pricing. the fact that this level of capability is becoming something you can just run.&lt;/p&gt;

&lt;p&gt;the gap between frontier and open models keeps shrinking. but the gap between "technically available" and "actually usable" is closing faster. step 3.7 flash is a good model. but the ecosystem support, the quantization options, the clear deployment docs - that is what makes it useful.&lt;/p&gt;

&lt;p&gt;i still wonder about the terminal-bench gap though. if they close that in 3.9 flash, the conversation gets a lot more interesting.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>agents</category>
    </item>
    <item>
      <title>Anthropic Called Opus 4.8 Modest. The Internet Disagreed.</title>
      <dc:creator>Dishant Sharma</dc:creator>
      <pubDate>Sun, 04 Oct 2026 09:12:19 +0000</pubDate>
      <link>https://dev.to/dishant0406/anthropic-called-opus-48-modest-the-internet-disagreed-27eg</link>
      <guid>https://dev.to/dishant0406/anthropic-called-opus-48-modest-the-internet-disagreed-27eg</guid>
      <description>&lt;p&gt;Anthropic called Claude Opus 4.8 "a modest but tangible improvement" in their own launch post. That one line tells you more about where AI stands right now than any benchmark score on a spreadsheet.&lt;/p&gt;

&lt;p&gt;The model dropped on May 28, just 41 days after Opus 4.7. And 4.7 got a chilly reception. Users found it disappointing.&lt;/p&gt;

&lt;p&gt;TechCrunch noted the fast turnaround "may have something to do with the chilly reception to Opus 4.7." Competition is everywhere. OpenAI's Codex is gaining traction. Google's Gemini 3.5 Flash keeps shipping. Anthropic needed to put out something that felt like a real step.&lt;/p&gt;

&lt;p&gt;Simon Willison called it "refreshing" to see an AI lab honestly describe a release as a minor incremental improvement. And he is right. For once, a company making AI models said the quiet part out loud.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This is a small update. It does some things better. It is not going to change your life.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;But that has not stopped people from arguing about it across every corner of the internet.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Honesty as a feature&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here is the part that stood out most. Anthropic did not just talk about benchmarks. They focused on honesty instead.&lt;/p&gt;

&lt;p&gt;They trained the model to flag uncertainties and stop making unsupported claims. Their evaluations show Opus 4.8 is around four times less likely than its predecessor to let code flaws pass without a word.&lt;/p&gt;

&lt;p&gt;The system card puts it plainly: "Claude Opus 4.8 had the lowest incorrect-rate of the six models on every benchmark." It achieved this mainly by saying "i don't know" more often instead of guessing wrong.&lt;/p&gt;

&lt;p&gt;i find this genuinely interesting. We spend so much time complaining about AI hallucinating. But when a model finally gets better at staying quiet about things it isn't sure about, we barely notice.&lt;/p&gt;

&lt;p&gt;The Bridgewater team did notice. In their testimonial, they said the biggest difference was "Opus 4.8's tendency to proactively flag issues with the inputs and outputs of an analysis, something other models routinely missed and left to the users to catch."&lt;/p&gt;

&lt;p&gt;But here is the other side. Some people think this makes the model worse at their job. The r/ClaudeCode thread titled "Pack it up, boys. Opus 4.8 is officially dead" captures the frustration.&lt;/p&gt;

&lt;p&gt;More hesitation. More caution. More second-guessing. That is not what everyone wants from a coding assistant.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The model is better at saying no. Whether that is good or bad depends on what you ask it to do.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;What else actually changed&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There is more than just honesty. Opus 4.8 introduces effort controls on claude.ai and Cowork. You can dial how much thinking Claude puts into a response.&lt;/p&gt;

&lt;p&gt;Low effort means faster replies. Max effort means it thinks harder. You pick the balance.&lt;/p&gt;

&lt;p&gt;Dynamic Workflows launched in research preview for Claude Code. The model can run hundreds of parallel subagents in one session. It can handle codebase-scale migrations across hundreds of thousands of lines, then verify its outputs before reporting back. That sounds ambitious until you remember Devin's CEO Scott Wu confirmed it fixes the "comment-verbosity and tool-calling issues" they saw with Opus 4.7.&lt;/p&gt;

&lt;p&gt;Fast mode is three times cheaper now. $10 per million input tokens against $30 for previous versions. That matters if you are running agent loops at scale.&lt;/p&gt;

&lt;p&gt;The API now accepts system entries inside the messages array. You can update Claude's instructions mid-task without breaking the prompt cache. Simon Willison noted the lower cache minimum too: 1,024 tokens instead of 4,096.&lt;/p&gt;

&lt;p&gt;Small changes. But they add up when you build on top of them.&lt;/p&gt;

&lt;p&gt;Here is a quick comparison of what shifted:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Area&lt;/th&gt;
&lt;th&gt;Opus 4.7&lt;/th&gt;
&lt;th&gt;Opus 4.8&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;SWE-bench Pro&lt;/td&gt;
&lt;td&gt;67.1%&lt;/td&gt;
&lt;td&gt;69.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code flaw detection&lt;/td&gt;
&lt;td&gt;Standard&lt;/td&gt;
&lt;td&gt;4x better&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fast mode cost&lt;/td&gt;
&lt;td&gt;$30/$150&lt;/td&gt;
&lt;td&gt;$10/$50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt cache min&lt;/td&gt;
&lt;td&gt;4,096 tokens&lt;/td&gt;
&lt;td&gt;1,024 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Honesty&lt;/td&gt;
&lt;td&gt;Guess first&lt;/td&gt;
&lt;td&gt;Flag first&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;The split in opinion&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The Reddit threads tell their own story. The official release post on r/ClaudeAI has 2.6K upvotes and 784 comments. Real interest is there.&lt;/p&gt;

&lt;p&gt;But head over to r/codex and you will find a thread titled "Opus 4.8 is not a step forward. It's Anthropic finally catching up." The argument goes that Claude is better at design, but Codex is better at coding.&lt;/p&gt;

&lt;p&gt;And that is the honest tension right now. Opus 4.8 beats previous Opus models on almost every benchmark. Cursor's CEO said it "exceeds prior Opus models across every effort level" on CursorBench. Databricks called it a real gain in agentic reasoning.&lt;/p&gt;

&lt;p&gt;But the bar has moved. Codex exists. Gemini 3.5 Flash exists. The question is no longer "is it better than Opus 4.7" but "is it the best available today."&lt;/p&gt;

&lt;p&gt;The answer depends entirely on what you are doing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The time i asked for honesty and got it&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;i spent an evening testing Opus 4.8 on a migration task. A small Django app, maybe 12 files. i asked it to refactor the models layer and port the views over. Old Claude would have charged ahead with confidence and left me debugging two broken endpoints at midnight.&lt;/p&gt;

&lt;p&gt;Opus 4.8 stopped three times to ask clarifying questions. Annoying? Yes. But the migration worked on the first try.&lt;/p&gt;

&lt;p&gt;i do not know if that is the kind of thing benchmarks capture. But it is the kind of thing that makes you trust the tool more.&lt;/p&gt;

&lt;p&gt;It reminded me of that coworker who triple-checks everything before merging. Frustrating to work with. Great to deploy after.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who actually needs this one&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Let me be direct. Most people do not need Opus 4.8.&lt;/p&gt;

&lt;p&gt;If you use Claude for chat, brainstorming, or writing drafts, the differences between 4.7 and 4.8 are barely noticeable. The model is more honest about what it does not know. That might actually make it less satisfying for casual conversation. You might feel like it is hesitating more.&lt;/p&gt;

&lt;p&gt;If you run agentic workloads, the improvements matter. Coding assistants, automated research, multi-step analysis. The honesty thing pays off when your agent makes decisions autonomously.&lt;/p&gt;

&lt;p&gt;The dynamic workflows feature is genuinely useful if you do codebase-scale work. The cost cuts on fast mode matter at volume.&lt;/p&gt;

&lt;p&gt;But for the average chat user? This is a maintenance release with an honest label. Nothing wrong with that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One last thought&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Anthropic ended their launch post with something unexpected. They said there is "still more to be done" and mentioned Project Glasswing. A model class with "even higher intelligence than Opus," currently in limited preview for cybersecurity work. They expect to bring it to everyone "in the coming weeks."&lt;/p&gt;

&lt;p&gt;That is the part i keep thinking about. While everyone is busy litigating whether Opus 4.8 is good enough on Reddit, Anthropic is already testing a model they will not even release widely yet because the safety work is not done. The real next step is sitting behind a door we cannot open.&lt;/p&gt;

&lt;p&gt;i think that is the interesting story here. Not whether 4.8 is a step forward. It is, modestly and tangibly. But in an industry that will not stop promising revolutions, the company that said "this is a small improvement" might be the one with the most interesting thing coming next.&lt;/p&gt;

</description>
      <category>claude</category>
    </item>
    <item>
      <title>The hype around GPT-Realtime-2 and what actually landed</title>
      <dc:creator>Dishant Sharma</dc:creator>
      <pubDate>Sun, 04 Oct 2026 09:11:46 +0000</pubDate>
      <link>https://dev.to/dishant0406/the-hype-around-gpt-realtime-2-and-what-actually-landed-nag</link>
      <guid>https://dev.to/dishant0406/the-hype-around-gpt-realtime-2-and-what-actually-landed-nag</guid>
      <description>&lt;p&gt;Thirty four points and five Hacker News comments is not what a voice breakthrough is supposed to look like. That was one of the first weird signals around OpenAI’s May 7 release of &lt;code&gt;GPT-Realtime-2&lt;/code&gt;, &lt;code&gt;GPT-Realtime-Translate&lt;/code&gt;, and &lt;code&gt;GPT-Realtime-Whisper&lt;/code&gt;. On paper, it looked loud.&lt;/p&gt;

&lt;p&gt;In the room, it felt quieter.&lt;/p&gt;

&lt;p&gt;OpenAI said the new lineup pushes realtime audio past simple call and response. The headline line was easy to repeat: GPT-5-class reasoning, live translation, live transcription, and more natural voice agents. And yeah, that sounds big.&lt;/p&gt;

&lt;p&gt;But the first real reaction i saw was less "holy shit" and more people squinting at what had actually shipped.&lt;/p&gt;

&lt;p&gt;That matters because voice AI has been stuck in the same awkward zone for a while. Demos sound smooth. Real apps still stumble when the user interrupts, changes direction, or asks for something that needs tools and memory at the same time.&lt;/p&gt;

&lt;p&gt;You’ve probably seen this. i know i have.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The hype was real. The instant takeover was not.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  What got people excited
&lt;/h3&gt;

&lt;p&gt;i used to think the main story here would be the translation model. It was not. The thing people kept repeating was &lt;code&gt;GPT-Realtime-2&lt;/code&gt; being OpenAI’s first voice model with GPT-5-class reasoning.&lt;/p&gt;

&lt;p&gt;That phrase does a lot of work.&lt;/p&gt;

&lt;p&gt;OpenAI also tucked in details that matter more than the branding:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;context jumped from &lt;code&gt;32K&lt;/code&gt; to &lt;code&gt;128K&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;reasoning can be set from &lt;code&gt;minimal&lt;/code&gt; to &lt;code&gt;xhigh&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;low&lt;/code&gt; is the default&lt;/li&gt;
&lt;li&gt;the model can use short preambles while it works&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last bit sounds small. It is not small.&lt;/p&gt;

&lt;p&gt;If you have ever used a voice agent that goes silent for two seconds, you know the problem. Users think it froze. OpenAI is basically productizing the fake little human noises support reps use so people do not hang up.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;voice AI is not just about sounding human. it is about not feeling broken.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The translation side is also bigger than it first looks. OpenAI says &lt;code&gt;GPT-Realtime-Translate&lt;/code&gt; can take speech from 70+ input languages into 13 output languages while keeping pace with the speaker. That is less sexy than a benchmark chart.&lt;/p&gt;

&lt;p&gt;But for actual travel, support, and operations apps, it is easier to sell than "smarter voice."&lt;/p&gt;

&lt;h3&gt;
  
  
  What actually shipped vs what people imagined
&lt;/h3&gt;

&lt;p&gt;what actually happens is people hear "GPT-5-class reasoning" and mentally jump to a sci-fi operator that can run your life. Then they open the release and find a very developer-shaped product.&lt;/p&gt;

&lt;p&gt;This launch was not really for casual users. It was for teams building call flows, support tools, booking assistants, and bilingual live experiences. OpenAI even framed it around patterns like voice-to-action, systems-to-voice, and voice-to-voice.&lt;/p&gt;

&lt;p&gt;Zillow, Deutsche Telekom, and Priceline got named because OpenAI wanted everyone to picture business workflows, not just talking avatars.&lt;/p&gt;

&lt;p&gt;Here’s the split i kept seeing:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;hype version&lt;/th&gt;
&lt;th&gt;shipped version&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;finally, AGI but on the phone&lt;/td&gt;
&lt;td&gt;better infrastructure for voice agents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;human-like chat in real time&lt;/td&gt;
&lt;td&gt;more reliable tool use and recovery&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;magic conversations&lt;/td&gt;
&lt;td&gt;product teams tuning latency, tone, and context&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;But that is not a knock.&lt;/p&gt;

&lt;p&gt;It is the part most people miss. Real upgrades in AI often look boring at launch because they fix the layer below the visible layer. A bigger context window, clearer recovery behavior, and audible tool transparency will not trend the same way as a wild demo clip.&lt;/p&gt;

&lt;p&gt;They do matter more once money is on the line.&lt;/p&gt;

&lt;h3&gt;
  
  
  The underreported detail
&lt;/h3&gt;

&lt;p&gt;here's a question people always ask: what was the surprising part nobody posted about enough?&lt;/p&gt;

&lt;p&gt;For me, it was the reasoning control. OpenAI did not just ship a smarter voice model. It shipped knobs.&lt;/p&gt;

&lt;p&gt;Developers can pick &lt;code&gt;minimal&lt;/code&gt;, &lt;code&gt;low&lt;/code&gt;, &lt;code&gt;medium&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt;, or &lt;code&gt;xhigh&lt;/code&gt; reasoning, with &lt;code&gt;low&lt;/code&gt; as default.&lt;/p&gt;

&lt;p&gt;That tells you something honest about the current state of voice AI. Everyone says they want a genius voice agent. Most people actually want a fast one that does not screw up basic tasks.&lt;/p&gt;

&lt;p&gt;And OpenAI knows it. If &lt;code&gt;low&lt;/code&gt; is default, that means latency still rules the room.&lt;/p&gt;

&lt;p&gt;There was another small tell. OpenAI highlighted recovery behavior almost as much as intelligence. That means failure is still central to the product story.&lt;/p&gt;

&lt;p&gt;The model is better because it can say "i'm having trouble with that right now" instead of silently dying in the middle of a workflow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That is progress. It is also an admission.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And yes, the launch page itself had a time-limited demo. Even the company selling the future still puts guardrails around how long you can poke the thing.&lt;/p&gt;




&lt;h3&gt;
  
  
  The reaction gap
&lt;/h3&gt;

&lt;p&gt;the first time i tried to gauge the public pulse on launches like this, i made the mistake of only reading company posts and benchmark summaries. Everything looked massive. Then i checked Reddit and HN, and the tone got more useful.&lt;/p&gt;

&lt;p&gt;The Reddit chatter, at least from surfaced forum results, leaned predictable: "smartest voice model yet," "new generation," "this could disrupt X." You know the script.&lt;/p&gt;

&lt;p&gt;Fast excitement. A lot of people repeating the release notes back to each other.&lt;/p&gt;

&lt;p&gt;Hacker News felt flatter. The submission for OpenAI’s voice release had real attention, but not the kind that makes you think everyone just dropped their current stack. That mismatch is useful.&lt;/p&gt;

&lt;p&gt;It says the launch hit the "interesting" bucket before the "must switch now" bucket.&lt;/p&gt;

&lt;p&gt;Here’s what that usually means.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;people believe the direction&lt;/li&gt;
&lt;li&gt;they do not yet trust the operational reality&lt;/li&gt;
&lt;li&gt;they are waiting for someone else to eat the bugs first&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;But i get it. Voice apps have burned a lot of goodwill.&lt;/p&gt;

&lt;p&gt;A slick demo is easy. A tool-calling voice agent that survives interruptions, wrong assumptions, accents, background noise, and messy human phrasing is still annoying to build. The hype is not fake.&lt;/p&gt;

&lt;p&gt;It is just ahead of the proof.&lt;/p&gt;

&lt;h3&gt;
  
  
  Small tangent, but it matters
&lt;/h3&gt;

&lt;p&gt;Most AI model names sound like someone lost a bet in a meeting room. &lt;code&gt;GPT-Realtime-2&lt;/code&gt; is not awful, but it still has that sterile product smell. i miss when software names had a little chaos.&lt;/p&gt;

&lt;p&gt;And voice products make this worse. The stuff is supposed to feel warm and immediate, but the names sound like firmware updates. Imagine explaining to your friend that your startup got saved by &lt;code&gt;GPT-Realtime-Whisper&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It sounds less like a breakthrough and more like a very anxious Bluetooth speaker.&lt;/p&gt;

&lt;p&gt;But naming matters because it shapes expectation. If you call something a realtime model with GPT-5-class reasoning, people will expect a near-human operator. If what they actually get is a steadier stack for support automation, they will call it overhyped even if the product is genuinely better.&lt;/p&gt;

&lt;p&gt;That gap between label and lived experience causes half the backlash in AI.&lt;/p&gt;




&lt;h3&gt;
  
  
  The real talk
&lt;/h3&gt;

&lt;p&gt;Most people do not need this.&lt;/p&gt;

&lt;p&gt;If you are a solo builder making a simple app, you probably should not sprint into realtime voice just because OpenAI launched a shinier stack. Voice still adds cost, latency tuning, prompt weirdness, tool orchestration, fallback design, and support headaches.&lt;/p&gt;

&lt;p&gt;It is one of the easiest ways to make a product feel impressive in a demo and exhausting in production.&lt;/p&gt;

&lt;p&gt;And if your use case works fine with text, stick to text. Seriously.&lt;/p&gt;

&lt;p&gt;The biggest misconception about this launch is that smarter voice means voice is suddenly the best interface for everything. It does not. It means the ceiling moved up for cases where hands-free use, live translation, or conversational task flow already made sense.&lt;/p&gt;

&lt;p&gt;For everyone else, this is mostly infrastructure news.&lt;/p&gt;

&lt;p&gt;That is not a bad thing. Infrastructure news becomes product reality later.&lt;/p&gt;

&lt;p&gt;But you should know which layer you are looking at before you start tweeting like the phone just got reinvented.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where i land on it
&lt;/h3&gt;

&lt;p&gt;i keep coming back to that quiet reaction on launch day. Big promise. Useful release.&lt;/p&gt;

&lt;p&gt;Smaller blast radius than the marketing language suggested.&lt;/p&gt;

&lt;p&gt;But i would still bet this launch matters.&lt;/p&gt;

&lt;p&gt;Not because people will remember the exact model names. They will not. It matters because OpenAI is getting more explicit about the ugly parts of voice products: recovery, latency, interruptions, tools, pacing, context.&lt;/p&gt;

&lt;p&gt;That is where real adoption lives.&lt;/p&gt;

&lt;p&gt;So no, i do not think May 7 was the day voice AI fully arrived. i think it was the day the sales pitch got a little closer to the work. And honestly, that is more interesting.&lt;/p&gt;

</description>
      <category>openai</category>
    </item>
    <item>
      <title>Why Developers Are Switching From Claude Code to Codex</title>
      <dc:creator>Dishant Sharma</dc:creator>
      <pubDate>Sun, 04 Oct 2026 09:11:44 +0000</pubDate>
      <link>https://dev.to/dishant0406/why-developers-are-switching-from-claude-code-to-codex-3ba0</link>
      <guid>https://dev.to/dishant0406/why-developers-are-switching-from-claude-code-to-codex-3ba0</guid>
      <description>&lt;p&gt;"I’ve switched almost all my usage to codex. For the simple reason that it’s more reliable." That line showed up in an Ask HN thread after another wave of Codex chatter hit Reddit, and it says more than most launch posts do.&lt;/p&gt;

&lt;p&gt;People are not switching because of one benchmark. They are switching because they are tired.&lt;/p&gt;

&lt;p&gt;Tired of burning a five-hour window on a half-finished feature. Tired of getting flashy output that still needs three review passes. Tired of arguing with a tool that was supposed to save time.&lt;/p&gt;

&lt;h2&gt;
  
  
  what people are actually reacting to
&lt;/h2&gt;

&lt;p&gt;OpenAI gave the hype machine plenty to work with. On April 16, 2026, it said more than &lt;strong&gt;3 million developers&lt;/strong&gt; were already using Codex every week, then added a long list of new tricks: background computer use, an in-app browser, memory, recurring automations, SSH into remote devboxes, PR review, and 90-plus plugins.&lt;/p&gt;

&lt;p&gt;That is plenty.&lt;/p&gt;

&lt;p&gt;Claude Code still has its own pull. It is easier to describe. Open the terminal, point it at a repo, let it read, edit, run commands, and work through the mess with you. You can feel that in the docs. Claude Code still sounds like a coding tool first. Codex now sounds like a whole workbench.&lt;/p&gt;

&lt;p&gt;And that difference matters because the switch story is not really about model IQ. It is about workflow gravity. If one tool starts feeling like the place where planning, review, browser checks, and long-running work all happen, people start moving there even before they fully trust it.&lt;/p&gt;




&lt;h2&gt;
  
  
  the real reasons
&lt;/h2&gt;

&lt;p&gt;here's a question people always ask: if Claude Code is still good, why are people even talking about switching?&lt;/p&gt;

&lt;p&gt;Because the reactions are weirdly consistent. On Reddit, people who use both keep saying some version of the same thing. Claude is better for UI and fast implementation. Codex is slower, more conservative, and much better at review. One person put it in plain English: Claude is the kid trying to impress you. Codex is the senior who moves slower and misses less.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That is the whole mood right now.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And hype loves a clean character split. One tool feels fast. One feels safe. Then the posts write themselves.&lt;/p&gt;

&lt;h3&gt;
  
  
  review beats vibes
&lt;/h3&gt;

&lt;p&gt;i used to think the hype was just OpenAI doing OpenAI things. Big launch. Big claims. Big social spillover.&lt;/p&gt;

&lt;p&gt;But the Reddit thread that caught my eye was not fanboy stuff. It was practical. People were saying they use Claude for drafting, then hand the work to Codex for review because Codex catches second-order issues, asks for fewer repair passes, and is easier to trust on large pull requests.&lt;/p&gt;

&lt;p&gt;That is not glamorous. It is not a real switch yet.&lt;/p&gt;

&lt;p&gt;It is a wedge. And wedges matter more than slogans. If a tool becomes your final reviewer, it slowly becomes your bossy friend in the room. Then one week later it becomes the tool you open first.&lt;/p&gt;

&lt;p&gt;Here’s what people keep describing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude gets you moving fast&lt;/li&gt;
&lt;li&gt;Codex slows you down at the right moment&lt;/li&gt;
&lt;li&gt;the combo often beats picking one side&lt;/li&gt;
&lt;li&gt;the switch usually starts in review, not generation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You have probably done this with coworkers too.&lt;/p&gt;

&lt;h3&gt;
  
  
  the limits story
&lt;/h3&gt;

&lt;p&gt;most tutorials tell you people switch because one model is smarter. i don't buy that.&lt;/p&gt;

&lt;p&gt;A lot of this is quota psychology. In that same Reddit discussion, people talked about timing Claude’s usage windows, burning through Sonnet, and treating Codex as the extra runway that keeps the day from stopping. On HN, one comment laid out a whole relay race: start with Claude, implement with Claude, hit the limit, switch to Codex for review, then bounce back.&lt;/p&gt;

&lt;p&gt;That is not devotion. That is coping.&lt;/p&gt;

&lt;p&gt;i have done the dumb thing of trusting a second green check and regretting it.&lt;/p&gt;

&lt;p&gt;And products get mistaken for movements when users are routing around friction. If your day depends on not losing momentum, the tool with looser limits or steadier behavior starts to look smarter than it is.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;most “switching” stories are really “my main tool annoyed me at the wrong time” stories.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But that is still real market behavior. Annoyance compounds.&lt;/p&gt;

&lt;h3&gt;
  
  
  the bigger shell
&lt;/h3&gt;

&lt;p&gt;what actually happens is people stop comparing model answers and start comparing surfaces.&lt;/p&gt;

&lt;p&gt;Claude Code still feels sharp inside the terminal. That matters. It is clean. It is direct. You say what you want, it inspects files, runs commands, and gets to work. But OpenAI is pushing a wider frame. Codex is becoming the app where code review, browser iteration, image work, automations, SSH sessions, and memory all sit in one place.&lt;/p&gt;

&lt;p&gt;That broader shell creates hype even for people who do not need half the features. Why? Because it sounds like leverage on future work. You may not need recurring automations today. But you can picture needing them next month. And product hype feeds on that future self.&lt;/p&gt;

&lt;p&gt;Here’s the simple version:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;thing&lt;/th&gt;
&lt;th&gt;Claude Code&lt;/th&gt;
&lt;th&gt;Codex&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;core feel&lt;/td&gt;
&lt;td&gt;terminal-first&lt;/td&gt;
&lt;td&gt;workspace-first&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;crowd story&lt;/td&gt;
&lt;td&gt;faster maker&lt;/td&gt;
&lt;td&gt;slower reviewer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;switch trigger&lt;/td&gt;
&lt;td&gt;limits, drift&lt;/td&gt;
&lt;td&gt;trust, breadth&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;But broader is not always better. Bigger surfaces also create more places for the tool to get weird.&lt;/p&gt;

&lt;h3&gt;
  
  
  the awkward part nobody likes
&lt;/h3&gt;

&lt;p&gt;the problem isn't what you think.&lt;/p&gt;

&lt;p&gt;The biggest underreported detail is that while people were hyping these coding agents, security researchers were finding ways to break them. VentureBeat reported that Claude Code, Copilot, and Codex all got hit by real exploit chains, and the pattern was boring in the worst way: attackers went after credentials, not model magic.&lt;/p&gt;

&lt;p&gt;Codex had a branch-name exploit that could expose a GitHub OAuth token. Claude Code had multiple permission and sandbox issues, including a deny-rule failure once commands crossed a certain length. None of that kills the tools. But it does puncture the fantasy that one of them is some calm senior engineer living inside your laptop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;These tools are still messy software.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And the people switching from Claude to Codex are not leaving one clean reality for another. They are picking a different mess that currently fits their nerves better.&lt;/p&gt;




&lt;h2&gt;
  
  
  the side quest
&lt;/h2&gt;

&lt;p&gt;This whole Claude-versus-Codex thing also reminds me how much developers love naming tools like they are fantasy weapons or old notebooks found in an attic. Claude Code sounds like a polite person in a sweater who shows up early. Codex sounds like the dangerous book you were not supposed to open.&lt;/p&gt;

&lt;p&gt;That matters more than people admit.&lt;/p&gt;

&lt;p&gt;Names leak into expectations. If a product is called Codex, you forgive a little drama because drama was already in the box. If a product is called Claude Code, you expect calm competence and maybe a slight sigh when it sees your repo. And when the tool fails, the name somehow makes the failure feel personal.&lt;/p&gt;

&lt;p&gt;i know that sounds dumb. But product markets are full of tiny emotional cheats like this.&lt;/p&gt;

&lt;p&gt;One reason the hype spreads so fast is that people are not only testing software. They are trying on identities. The serious reviewer. The speed demon. The person who uses both because they are above the war.&lt;/p&gt;

&lt;h2&gt;
  
  
  the blunt version
&lt;/h2&gt;

&lt;p&gt;Most people do not need to switch.&lt;/p&gt;

&lt;p&gt;If Claude Code already fits your repo, your habits, and your patience, keep using it. If Codex review catches issues your current setup misses, add Codex for review and stop pretending you need a religious conversion. A lot of the loudest posts are workflow cosplay wrapped around tool choice.&lt;/p&gt;

&lt;p&gt;But some people should switch, or at least test hard.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;If review quality matters more than raw speed, try Codex.&lt;/li&gt;
&lt;li&gt;If you keep hitting Claude limits at bad moments, try Codex.&lt;/li&gt;
&lt;li&gt;If you want one wider workspace instead of a tighter terminal tool, try Codex.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Do not switch because the internet got loud. Switch because one repeated pain keeps costing you real hours.&lt;/p&gt;

&lt;p&gt;And if you are early in your career, be careful. The easy trap is letting tool preference replace actual engineering judgment. A different harness will not save you from vague specs, weak tests, or bad taste.&lt;/p&gt;

&lt;h2&gt;
  
  
  where this probably lands
&lt;/h2&gt;

&lt;p&gt;i don't think the story ends with people fully leaving Claude Code. i think it lands somewhere more annoying.&lt;/p&gt;

&lt;p&gt;A lot of builders will keep two windows open. Claude for momentum. Codex for review. Claude for the rough draft. Codex for the second look. Then another tool will show up and promise to beat both by combining the good parts and dropping weird parts.&lt;/p&gt;

&lt;p&gt;But the recent hype does reveal something useful. People are done clapping for demos. They want tools that feel steady at hour six, not exciting at minute six.&lt;/p&gt;

&lt;p&gt;And that is why the switch talk matters. Not because everyone is leaving Claude Code. Because reliability has finally become more marketable than speed.&lt;/p&gt;

&lt;p&gt;That means the category is growing up.&lt;/p&gt;

</description>
      <category>claude</category>
    </item>
    <item>
      <title>GPT-5.5 and Codex: What the Hype Missed and What Actually Matters</title>
      <dc:creator>Dishant Sharma</dc:creator>
      <pubDate>Sun, 04 Oct 2026 09:11:11 +0000</pubDate>
      <link>https://dev.to/dishant0406/gpt-55-and-codex-what-the-hype-missed-and-what-actually-matters-3jfh</link>
      <guid>https://dev.to/dishant0406/gpt-55-and-codex-what-the-hype-missed-and-what-actually-matters-3jfh</guid>
      <description>&lt;p&gt;An NVIDIA engineer who tested GPT-5.5 early access said losing it felt like losing a limb. Dan Shipper, CEO of Every, called it the first coding model with "serious conceptual clarity." Pietro Schirano watched it merge hundreds of frontend changes into a main branch that had also shifted substantially, finishing in one shot in about 20 minutes. These are not casual impressions. These are people who build with AI daily, and they sound genuinely surprised.&lt;/p&gt;

&lt;p&gt;GPT-5.5 dropped on April 23, 2026. Codex got a massive desktop update a week earlier on April 16. OpenAI dropped both in the same month, and the internet lost its mind for about 48 hours. Then the complaints started.&lt;/p&gt;




&lt;h2&gt;
  
  
  What GPT-5.5 actually is
&lt;/h2&gt;

&lt;p&gt;Here is the short version. GPT-5.5 is the first fully retrained base model OpenAI has shipped since GPT-4.5. Everything between 4.5 and 5.5 was incremental. This one is a new foundation.&lt;/p&gt;

&lt;p&gt;The benchmarks look strong. 82.7% on Terminal-Bench 2.0. 78.7% on OSWorld-Verified. 58.6% on SWE-Bench Pro. 73.1% on Expert-SWE, an internal eval where the median human completion time is 20 hours.&lt;/p&gt;

&lt;p&gt;But benchmarks lie, so here is what matters more. GPT-5.5 uses &lt;strong&gt;40% fewer tokens&lt;/strong&gt; than GPT-5.4 on the same Codex tasks. It matches GPT-5.4 per-token latency despite being a bigger, smarter model. That is not nothing. That is the kind of efficiency improvement that actually changes how you work with it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The model helped build the infrastructure that serves it.&lt;/strong&gt; OpenAI used Codex and GPT-5.5 to optimize their own inference stack on NVIDIA GB200 and GB300 systems. One result was custom load-balancing heuristics that boosted token generation speed by over 20%. That is a weird, recursive detail that most coverage skipped.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Codex desktop update
&lt;/h2&gt;

&lt;p&gt;A week before GPT-5.5, OpenAI shipped a major Codex update that barely got attention because everyone was waiting for the model. That update matters more than people realize.&lt;/p&gt;

&lt;p&gt;Here is what changed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Background computer use&lt;/strong&gt;: Codex agents can now see your screen, click, and type with their own cursor in the background. Multiple agents can work in parallel on your Mac without blocking your own work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;In-app browser&lt;/strong&gt;: You can comment directly on web pages to give Codex precise instructions. Built for frontend and game dev iteration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Image generation with &lt;code&gt;gpt-image-1.5&lt;/code&gt;&lt;/strong&gt;: Create mockups, game assets, and product visuals inside the same workflow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory preview&lt;/strong&gt;: Codex remembers preferences, corrections, and context from previous sessions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;90+ new plugins&lt;/strong&gt;: Atlassian Rovo, CircleCI, GitLab Issues, Microsoft Suite, Neon, Render, and more.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SSH remote devbox support&lt;/strong&gt;: Connect to remote machines directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automated scheduling&lt;/strong&gt;: Codex can schedule future work for itself and wake up automatically to continue long-running tasks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Codex desktop app launched February 2, 2026 for macOS. Windows followed March 4. This April 16 update is the first real expansion that makes it competitive with Claude Code and Cursor as a full-workflow tool.&lt;/p&gt;




&lt;h2&gt;
  
  
  The price problem
&lt;/h2&gt;

&lt;p&gt;$5 per million input tokens. $30 per million output tokens. That is exactly double what GPT-5.4 costs.&lt;/p&gt;

&lt;p&gt;Reddit noticed immediately. A thread on r/theprimeagen with 40+ comments called out the doubling. Someone on r/codex asked if the improvements justify the price bump. The top reply was blunt: "that's why I still use gpt-5.3-codex as my backbone model. best performance over cost so far."&lt;/p&gt;

&lt;p&gt;OpenAI's argument is that the 40% token reduction offsets the higher per-token price. That math works if you actually use fewer tokens. It does not work if your workload expands to fill whatever budget the model gives you, which is what usually happens with better tools.&lt;/p&gt;

&lt;p&gt;GPT-5.5 Pro is even steeper: $30 input, $180 output. That is not a typo.&lt;/p&gt;




&lt;h2&gt;
  
  
  What the internet gets wrong about model releases
&lt;/h2&gt;

&lt;p&gt;Every time a frontier model drops, the same cycle plays out. Leaks surface. Codenames get romanticized. People convince themselves the next one will change everything overnight. Then it ships, and the hot take splits into "this is everything" and "this is nothing."&lt;/p&gt;

&lt;p&gt;It happened with GPT-5.2. That model launched to similar hype. Four months later it was sitting at number 15 on the LMArena leaderboard, behind GPT-5.1 and Claude 4.5. The people who built real products with it did not care about the ranking. The people who were watching the ranking moved on to the next one.&lt;/p&gt;

&lt;p&gt;GPT-5.5 will probably follow the same arc. The real question is not whether it beats Claude Opus 4.7 on Terminal-Bench by 0.7 points. The real question is whether the Codex desktop app, with computer use and memory and background agents, changes how a developer spends their morning.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Spud backlash
&lt;/h2&gt;

&lt;p&gt;Internally, GPT-5.5 was codenamed "Spud." The hype around that codename got absurd. People on r/accelerate compared it to Anthropic's "Mythos" and treated it like it would be GPT-6.&lt;/p&gt;

&lt;p&gt;Then it launched, and the reaction split in half.&lt;/p&gt;

&lt;p&gt;On r/singularity, a post titled "Big model feel with GPT 5.5" got 70+ comments. The top sentiment was genuine surprise at how much better it felt in practice. On r/vibecoding, someone said "i get the hype but opus still has something that isn't reflected on the benchmarks."&lt;/p&gt;

&lt;p&gt;The most honest take came from someone on r/accelerate: "this was hyped as if it was gpt 6, not a 0.1 improvement. from many people. please keep yourselves to a higher standard than that."&lt;/p&gt;

&lt;p&gt;That thread had 20+ comments and was posted just 21 hours after launch.&lt;/p&gt;




&lt;h2&gt;
  
  
  A math professor built an algebraic geometry app in 11 minutes
&lt;/h2&gt;

&lt;p&gt;Bartosz Naskrecki, an assistant professor of mathematics in Poland, gave Codex a single prompt to build an algebraic geometry app. The app visualizes quadratic surface intersections and converts the resulting curves into Weierstrass models using the computational Riemann-Roch theorem.&lt;/p&gt;

&lt;p&gt;It worked. Eleven minutes. He extended it with stable singularity visualization and exact reusable coefficients.&lt;/p&gt;

&lt;p&gt;He said the bigger shift is not any single app. It is that Codex can now help implement custom mathematical visualization and computer-algebra workflows that previously required dedicated tools. That is a specific, nerdy detail. But it is the kind of thing that shows where this is actually heading.&lt;/p&gt;

&lt;p&gt;OpenAI also claims an internal version of GPT-5.5 found a new proof about off-diagonal Ramsey numbers, later verified in Lean. Ramsey numbers are one of the central objects in combinatorics. Results in this area are rare and technically difficult. If that holds up, it is a genuine research contribution, not just a coding trick.&lt;/p&gt;




&lt;h2&gt;
  
  
  The real picture
&lt;/h2&gt;

&lt;p&gt;GPT-5.5 is good. Probably the best model you can use right now for agentic coding and complex task execution. But the gap between the Spud hype and the actual product annoyed people who expected a paradigm shift.&lt;/p&gt;

&lt;p&gt;And the doubling of API prices at a time when Anthropic and Google are pushing hard on competitive pricing makes the value conversation more complicated than the benchmark charts suggest.&lt;/p&gt;

&lt;p&gt;If you are a ChatGPT Plus subscriber, you get it for the same $20/month. That is a good deal. If you are running it through an API at scale, you need to do real math on whether the token savings actually cover the price jump.&lt;/p&gt;

&lt;p&gt;The Codex desktop update is the quieter story. Computer use in the background, memory, automated scheduling. That is the infra that makes the model actually useful in a daily workflow. Without it, GPT-5.5 is just a smarter API endpoint.&lt;/p&gt;

&lt;p&gt;i am still thinking about that NVIDIA engineer's quote. "losing access to GPT-5.5 feels like i've had a limb amputated." Maybe he was exaggerating. Maybe not. But no one said anything close to that about GPT-5.4.&lt;/p&gt;

</description>
      <category>codex</category>
      <category>gpt55</category>
    </item>
  </channel>
</rss>
