<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: TheRabbitHole</title>
    <description>The latest articles on DEV Community by TheRabbitHole (@therabbithole).</description>
    <link>https://dev.to/therabbithole</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F2473271%2Fb0e05429-f845-47c1-a1dd-73b29a1aa186.jpeg</url>
      <title>DEV Community: TheRabbitHole</title>
      <link>https://dev.to/therabbithole</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/therabbithole"/>
    <language>en</language>
    <item>
      <title>The Biggest AI Stories Weren’t Features. They Were Dependency</title>
      <dc:creator>TheRabbitHole</dc:creator>
      <pubDate>Thu, 30 Jul 2026 08:37:06 +0000</pubDate>
      <link>https://dev.to/therabbithole/the-biggest-ai-stories-werent-features-they-were-dependency-42jf</link>
      <guid>https://dev.to/therabbithole/the-biggest-ai-stories-werent-features-they-were-dependency-42jf</guid>
      <description>&lt;p&gt;Between May and late July 2026, OpenAI and Anthropic shipped more than thirty user-facing products and features: new voice modes, health-record integrations, personal finance, a desktop redesign, rewritten memory systems, an enterprise agent product, and two frontier-model families apiece.&lt;/p&gt;

&lt;p&gt;Most of them arrived with an announcement, a demo, a burst of press coverage—and then disappeared from public conversation.&lt;/p&gt;

&lt;p&gt;I wanted to understand the gap between what companies announced and what users found worth discussing. Not what the press covered; the press covers almost everything. I looked instead at public conversations on Hacker News, Reddit’s AI communities, and X, then ranked launches by a deliberately simple question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Did people care enough to argue about it?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That question reorganizes the last three months.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-sentence version
&lt;/h2&gt;

&lt;p&gt;Users rarely talked about conventional product features. They talked about models, prices, outages, privacy, and who controls access.&lt;/p&gt;

&lt;p&gt;Every top-tier discussion in this dataset concerned a model release, a governance fight, or something breaking. The few exceptions are instructive: they either solved a specific infrastructure problem or gave users something they could create and share with one another.&lt;/p&gt;

&lt;h2&gt;
  
  
  A note on the numbers
&lt;/h2&gt;

&lt;p&gt;This is a study of &lt;strong&gt;public attention&lt;/strong&gt;, not adoption, retention, revenue, or product value. Quiet features can be heavily used, and loud controversies can involve a relatively small population.&lt;/p&gt;

&lt;p&gt;I reviewed English-language discussions posted between May 1 and July 28, 2026, on Hacker News, prominent AI-related subreddits, and X. The figures below are raw engagement counts from each platform, not a composite score; an HN point, Reddit upvote, X like, and comment are not equivalent units. Keyword searches also miss conversations that use unexpected names.&lt;/p&gt;

&lt;p&gt;The cutoff particularly disadvantages late-July launches, which had only days to accumulate attention. Treat the rankings as a map of visible conversation—not a definitive measure of what succeeded.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tier 1: The stories that consumed the conversation
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. The Fable 5 / Mythos 5 export-control saga
&lt;/h3&gt;

&lt;p&gt;Anthropic launched Claude Fable 5 on June 9, its first publicly available Mythos-class model. Three days later, the US Department of Commerce issued a directive forcing Anthropic to suspend access to both Fable 5 and Mythos 5. The controls were lifted on June 30.&lt;/p&gt;

&lt;p&gt;This was not primarily a product story. It was a sovereignty story, and it dwarfed almost everything else:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Hacker News thread&lt;/th&gt;
&lt;th&gt;Points&lt;/th&gt;
&lt;th&gt;Comments&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Statement on US government directive to suspend access&lt;/td&gt;
&lt;td&gt;3,158&lt;/td&gt;
&lt;td&gt;2,314&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5 launch&lt;/td&gt;
&lt;td&gt;2,626&lt;/td&gt;
&lt;td&gt;2,160&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Commerce has lifted export controls&lt;/td&gt;
&lt;td&gt;977&lt;/td&gt;
&lt;td&gt;692&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Feds freaked over Fable 5 after “fix this code”&lt;/td&gt;
&lt;td&gt;613&lt;/td&gt;
&lt;td&gt;361&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fable 5 is Back&lt;/td&gt;
&lt;td&gt;408&lt;/td&gt;
&lt;td&gt;419&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Reddit followed just as closely. “Fable 5 is coming back!” received 5,338 upvotes and 506 comments; “Access has been extended!” received 5,273 and 777; “Fable staying on Max” received 3,279 and 724.&lt;/p&gt;

&lt;p&gt;What made the episode detonate was not only the model’s capability. A government had reached into a commercial service and switched off a model on which people already depended.&lt;/p&gt;

&lt;p&gt;The discussion immediately moved beyond benchmarks and into arms-control law. One widely upvoted comment noted that the directives fell under the Arms Export Control Act, potentially treating model weights as technical data under ITAR. Another observed that the restrictions reportedly prevented Anthropic’s own foreign employees from accessing Mythos internally—a constraint likely to be commercially intolerable.&lt;/p&gt;

&lt;p&gt;A substantial faction blamed Anthropic’s own political strategy:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Would the US government have slapped Anthropic with this export control if Anthropic never fearmonger’ed about Mythos? I think the answer is very likely no. […] This is a failure of Anthropic’s politicking.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Across dozens of threads, the durable conclusion was the same: access you do not control is conditional access. One commenter put it plainly:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Open weights + deterministic orchestration feels like the only sane long-term bet.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For much of June, this argument crowded almost everything else out.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. GPT-5.6 Sol—and who gets to use it
&lt;/h3&gt;

&lt;p&gt;OpenAI previewed GPT-5.6 Sol on June 26 and released it on July 9. It was the company’s largest story of the quarter by a wide margin: roughly 5,000 Hacker News points across the five leading threads, while the largest announcement post on X drew 6,212 likes and 976,000 views.&lt;/p&gt;

&lt;p&gt;But the distribution of attention matters:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Hacker News thread&lt;/th&gt;
&lt;th&gt;Points&lt;/th&gt;
&lt;th&gt;Comments&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 launch&lt;/td&gt;
&lt;td&gt;1,561&lt;/td&gt;
&lt;td&gt;1,113&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;U.S. government will decide who gets to use GPT-5.6&lt;/td&gt;
&lt;td&gt;1,184&lt;/td&gt;
&lt;td&gt;1,240&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Previewing GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;1,139&lt;/td&gt;
&lt;td&gt;744&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 used a prompt to close a 30-year gap in convex optimization&lt;/td&gt;
&lt;td&gt;600&lt;/td&gt;
&lt;td&gt;391&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture&lt;/td&gt;
&lt;td&gt;538&lt;/td&gt;
&lt;td&gt;443&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The access-control thread attracted more comments than the launch itself. It expressed the same anxiety as the Fable 5 episode, this time around a different vendor: people were evaluating not only what the model could do, but whether they could build on it without a third party later changing the terms.&lt;/p&gt;

&lt;p&gt;The other major branch of the Sol conversation was capability-real. Users discussed claimed mathematical advances and verified results rather than benchmark deltas. But practitioner complaints were equally concrete, especially when safety refusals consumed paid sessions:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“GPT-5.6 Sol finds a vulnerability but refuses to explain it. I think it would be a good practice to refund the session cost in that case. Otherwise a customer just spent some money in order to get exactly nothing.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The model earned attention through capability. The service around it earned scrutiny through control and pricing.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Claude Opus 5—and an immediate split
&lt;/h3&gt;

&lt;p&gt;Opus 5, released July 24, produced the largest single engagement number in the dataset. Anthropic’s announcement on X reached 61,266 likes and 23 million views. The Hacker News launch thread drew 1,777 points and 1,329 comments; Reddit’s drew 2,916 upvotes and 672 comments.&lt;/p&gt;

&lt;p&gt;Within 72 hours, however, the mood split:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;“Claude Opus 5 is ridiculously good at web design”—643 likes&lt;/li&gt;
&lt;li&gt;“How are we feeling about Opus 5 so far?”—2,453 likes&lt;/li&gt;
&lt;li&gt;“I do not like Opus 5 as much as I hoped to :(”—2,428 likes&lt;/li&gt;
&lt;li&gt;Three separate Hacker News threads about elevated Opus 5 errors reached the front page in four days&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The recurring technical complaint concerned defaults more than raw capability:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“I was using Opus 4.6 until 2 days ago with no CLAUDE.md or anything and it was great. Tried out Opus 5 and it’s been a super annoying experience out of the box.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A Reddit post titled “A week on Opus 5—best value at the frontier, but 3 default settings aren’t good” captured the emerging consensus. Users often liked the model and disliked the configuration in which it arrived.&lt;/p&gt;

&lt;p&gt;That distinction matters because it describes a fixable product problem. Benchmarks can reveal capability; forum complaints reveal the friction between that capability and its default presentation.&lt;/p&gt;

&lt;p&gt;Some heavy users still preferred the older, more expensive Fable 5 for difficult work:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“I believe Fable is the sharpest and most effective instrument I’ve ever used in AI.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  4. Claude Opus 4.8
&lt;/h3&gt;

&lt;p&gt;Opus 4.8, released May 28, generated 1,774 Hacker News points and 1,376 comments. Its central pitch—roughly four times less likely than 4.7 to overlook flaws in its own code—addressed a problem its audience had already been discussing.&lt;/p&gt;

&lt;p&gt;It produced one large thread, landed cleanly, and remained unusually uncontroversial. In this dataset, that is a compliment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tier 2: Real, but second-order
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Claude Code became a political object
&lt;/h3&gt;

&lt;p&gt;Claude Code appeared in 5,559 Hacker News comments, but its most prominent threads were rarely about new features:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Hacker News thread&lt;/th&gt;
&lt;th&gt;Points&lt;/th&gt;
&lt;th&gt;Comments&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Code is steganographically marking requests&lt;/td&gt;
&lt;td&gt;2,445&lt;/td&gt;
&lt;td&gt;750&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Code uses Bun written in Rust now&lt;/td&gt;
&lt;td&gt;608&lt;/td&gt;
&lt;td&gt;849&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k&lt;/td&gt;
&lt;td&gt;706&lt;/td&gt;
&lt;td&gt;396&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Microsoft starts canceling Claude Code licenses&lt;/td&gt;
&lt;td&gt;493&lt;/td&gt;
&lt;td&gt;466&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;I used Claude Code to get a second opinion on my MRI&lt;/td&gt;
&lt;td&gt;566&lt;/td&gt;
&lt;td&gt;715&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Trust, token efficiency, procurement, and off-label medical use dominated. The period’s actual feature releases—dynamic workflows, nested subagents, and background code review—barely registered independently.&lt;/p&gt;

&lt;p&gt;Claude subagents attracted 237 comments, and the strongest-performing subagent thread was itself a complaint: “Claude Code has a hardcoded instruction telling Opus 5 not to use subagents.”&lt;/p&gt;

&lt;p&gt;The lesson for developer-tool companies is simple: at sufficient scale, the changelog stops being the story. The product’s behavior becomes the story.&lt;/p&gt;

&lt;h3&gt;
  
  
  Usage limits: the feature nobody shipped and everybody discussed
&lt;/h3&gt;

&lt;p&gt;Usage-limit discussions produced 223 Hacker News comments, while Reddit’s “Dear Anthropic, This Has to STOP.” drew 2,673 upvotes and 555 comments.&lt;/p&gt;

&lt;p&gt;Pricing and rate limits generated more sustained, emotionally intense discussion than voice, health, memory, personal finance, and the desktop redesign combined.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“I’m not in a position to drop $200/month […] for coding tasks Kimi 2.6 has been about the same as Sonnet in my experience.”&lt;/p&gt;

&lt;p&gt;“That’s the difference between ‘I don’t use it for anything serious because I constantly run into limits’ and…”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every limits thread doubled as a churn thread. They read like unsolicited exit interviews and repeatedly named the same alternatives: Kimi, DeepSeek V4, GLM, and local models.&lt;/p&gt;

&lt;p&gt;This may be the most commercially dangerous conversation surrounding either company. It is not a feature problem; it is the point at which pricing and reliability determine whether capability can become habit.&lt;/p&gt;

&lt;h3&gt;
  
  
  Security and privacy incidents outperformed launches on Reddit
&lt;/h3&gt;

&lt;p&gt;The largest substantive Reddit post in the dataset was not a launch. “You can view a lot of shared conversations via Google” drew 8,066 upvotes and 1,340 comments, while a companion post about the same apparent exposure received 4,221 upvotes.&lt;/p&gt;

&lt;p&gt;Together, those posts attracted more engagement than Opus 5, Sonnet 5, and Claude Skills combined.&lt;/p&gt;

&lt;p&gt;Incidents travel further than announcements because they transform an abstract risk into a personal one. A feature asks users to imagine value. A privacy failure asks them to imagine themselves as the victim.&lt;/p&gt;

&lt;h3&gt;
  
  
  Claude Sonnet 5
&lt;/h3&gt;

&lt;p&gt;Sonnet 5, released June 30, received 1,266 Hacker News points, 784 comments, and 2,760 Reddit upvotes. Its one-million-token context window mattered to practitioners building context-heavy systems.&lt;/p&gt;

&lt;p&gt;It also received the period’s most brutal one-line headline—“Sonnet 5 Is Dead in the Water”—which seems to reflect, at least in part, the extraordinary saturation created by the Fable 5 political drama surrounding its launch.&lt;/p&gt;

&lt;h3&gt;
  
  
  Exception one: MCP’s stateless transport
&lt;/h3&gt;

&lt;p&gt;The July 28 MCP revision attracted only 118 Hacker News points and 37 comments. By raw volume it was niche, but the responses came from practitioners who had already encountered the problem and were unusually consistent:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“I’d made the shift to HTTP/Stateless from MCP a few months ago. It’s the right thing to do IMHO. Reliability up, problems down.”&lt;/p&gt;

&lt;p&gt;“URL Elicitation works well if a human is driving the client. Unfortunately MCP client support is patchy but I expect that will change now the protocol is stateless.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is the first exception to the general pattern. A conventional feature generated modest attention but high-quality discussion because it solved an identifiable infrastructure problem.&lt;/p&gt;

&lt;p&gt;There was a meaningful countercurrent. Some developers argued that the protocol remained unnecessary:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“You can get better results with skills SKILL.md + linked .md files with curl commands inside… Just plain HTTP(S).”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Low volume does not necessarily mean low importance. In infrastructure, a few dozen comments from people who have already migrated can be more informative than thousands of launch-day reactions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Exception two: Claude Skills spread by themselves
&lt;/h3&gt;

&lt;p&gt;On Hacker News, “Claude skill” was nearly a rounding error: 34 comments. On Reddit, however, Skills produced one of the strongest organic signals in the dataset—and the energy came from users rather than Anthropic:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;“I made a Claude Code skill that turns a photo of your handwriting into an installable font”—3,437 upvotes&lt;/li&gt;
&lt;li&gt;“New: Teach Claude a skill”—2,712 upvotes&lt;/li&gt;
&lt;li&gt;“Whoever created the ADHD skill god bless you”—2,621 upvotes and 405 comments&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The last example is the tell. It is not applause for a launch. It is one user thanking another for something that changed their day.&lt;/p&gt;

&lt;p&gt;The smaller Hacker News discussion also described genuine diffusion into nontechnical teams:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“That’s how I see most of my less technical coworkers reason about using AI. ‘Is there a Claude skill for that?’ is a question I hear multiple times a week.”&lt;/p&gt;

&lt;p&gt;“I thought about MCP, but found that having it as a Claude skill is much simpler (since it can be installed as a plugin, and only depends on md files and also doesn’t need to run a server all the time).”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Skills spread because they are user-authored, legible, and shareable. Most features in this dataset are things a company does for—or to—its users. Skills are things users do for one another. That gives the discussion a different character: organic rather than reactive.&lt;/p&gt;

&lt;p&gt;There is a counter-signal worth respecting. Some power users reported Skills degrading performance on newer models: “saw my skills start causing degradation”; “removing skills like ‘superpowers’ reduces token consumption.”&lt;/p&gt;

&lt;p&gt;Skills have a bloat problem. But bloat is often the problem successful platforms acquire after people begin building on them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Tier 0: Shipped into near-silence
&lt;/h2&gt;

&lt;p&gt;The following products were staffed, built, and announced during the same period. Within the sources and searches used for this analysis, their visible public reaction was minimal:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Shipped&lt;/th&gt;
&lt;th&gt;Discussion found&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;ChatGPT Voice in Chat, Work, and Codex&lt;/td&gt;
&lt;td&gt;Jul 23&lt;/td&gt;
&lt;td&gt;No dedicated discussion distinguishable from unrelated voice threads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New ChatGPT desktop app, unifying Chat, Work, and Codex&lt;/td&gt;
&lt;td&gt;Jul 16&lt;/td&gt;
&lt;td&gt;2 points, 1 comment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Health in ChatGPT, including Apple Health and medical records&lt;/td&gt;
&lt;td&gt;Jul 23&lt;/td&gt;
&lt;td&gt;32 points / 54 comments, plus 8 / 7&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI Presence, enterprise voice-and-chat agents&lt;/td&gt;
&lt;td&gt;Jul 22&lt;/td&gt;
&lt;td&gt;64 points, 51 comments in one thread&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rebuilt ChatGPT memory&lt;/td&gt;
&lt;td&gt;Jun 4&lt;/td&gt;
&lt;td&gt;No dedicated thread above background noise&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Codex Remote general availability&lt;/td&gt;
&lt;td&gt;May/Jun&lt;/td&gt;
&lt;td&gt;2 stories, 3 points total&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ChatGPT personal finance / Plaid&lt;/td&gt;
&lt;td&gt;May 15&lt;/td&gt;
&lt;td&gt;7 points, 1 comment&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retiring group chats&lt;/td&gt;
&lt;td&gt;Jul 9&lt;/td&gt;
&lt;td&gt;2 points, 0 comments&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.2 / GPT-4.5 deprecation&lt;/td&gt;
&lt;td&gt;Jun 12 / 26&lt;/td&gt;
&lt;td&gt;9 points total&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude voice-mode expansion&lt;/td&gt;
&lt;td&gt;Jul 23&lt;/td&gt;
&lt;td&gt;No signal distinguishable from background noise&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The silence does not establish that these products went unused. It does show that they failed to become stories within the communities examined.&lt;/p&gt;

&lt;h3&gt;
  
  
  Voice is expensive silence
&lt;/h3&gt;

&lt;p&gt;Both companies shipped substantial voice work within a day of each other in late July. Neither generated meaningful discussion in this dataset.&lt;/p&gt;

&lt;p&gt;Voice demos beautifully. Among people who publicly discuss software, however, it has yet to produce a comparable culture of use, argument, or user-created artifacts. That gap—between demo appeal and visible practitioner enthusiasm—is worth investigating with actual usage data.&lt;/p&gt;

&lt;h3&gt;
  
  
  Health deserved attention and received suspicion
&lt;/h3&gt;

&lt;p&gt;Connecting medical records to a general-purpose chatbot may be the most consequential product OpenAI shipped all quarter. The leading thread drew 32 points.&lt;/p&gt;

&lt;p&gt;The little discussion that did appear was skeptical. Adjacent headlines framed the product as “ChatGPT wants access to your health records so it can be a better not-doctor,” alongside coverage of a lawsuit.&lt;/p&gt;

&lt;p&gt;For a launch this sensitive, silence combined with distrust is not neutral. It suggests that the company has not yet established the legitimacy required for users to evaluate the value proposition on its own terms.&lt;/p&gt;

&lt;h3&gt;
  
  
  The memory rewrite arrived after users had formed their verdict
&lt;/h3&gt;

&lt;p&gt;The memory redesign appears to represent substantial engineering: time-aware resynthesis, costs reportedly reduced by roughly five times, and broader availability for free users.&lt;/p&gt;

&lt;p&gt;Yet users mostly discussed memory to complain about it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“This thing is better disabled because it’s intrusive.”&lt;/p&gt;

&lt;p&gt;“I found Opus’s 4.8 memories largely lacking value. I disabled memory for the web UI.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Improving a system after users have disabled it creates a distribution problem. The company is no longer merely shipping a better feature; it must persuade users to revisit a decision they believe they have already settled.&lt;/p&gt;

&lt;h3&gt;
  
  
  Codex Remote may have a naming problem
&lt;/h3&gt;

&lt;p&gt;Starting a coding task from a phone is plainly useful, and people were independently building similar products. Hacker News featured projects titled “Zedra—Mobile control plane for AI coding agents” and “ShellTeam.”&lt;/p&gt;

&lt;p&gt;Yet Codex Remote itself produced only three points of visible discussion.&lt;/p&gt;

&lt;p&gt;The market appeared to recognize the job while overlooking OpenAI’s implementation of it. That can happen when a product name describes internal architecture rather than the moment of user value. “Remote” says where the task runs. It does not say: start work from your phone, let it continue elsewhere, and return to a finished result.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the conversation was really about
&lt;/h2&gt;

&lt;p&gt;The public AI conversation of the last three months followed a surprisingly consistent hierarchy:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Capability earns attention.&lt;/strong&gt; Frontier models still create the largest launch moments, especially when they demonstrate results that feel real rather than benchmark-shaped.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Control turns attention into argument.&lt;/strong&gt; Export restrictions, procurement decisions, safety refusals, privacy incidents, and changing access terms dominate once users begin depending on a system.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Price and reliability determine habit.&lt;/strong&gt; A brilliant model that is unavailable, error-prone, or exhausted after a few sessions cannot become infrastructure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;User-created artifacts generate affection.&lt;/strong&gt; Skills stood apart because users could make them, exchange them, and thank one another for them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quiet features remain genuinely hard to judge from public discussion.&lt;/strong&gt; Silence may mean indifference, poor positioning, invisible success, or simply that a product does not produce stories.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That last point is the limit of this analysis, but it is also a useful product question. If a feature is valuable yet nobody talks about it, a company should know whether it has built quiet infrastructure or merely shipped into a void.&lt;/p&gt;

&lt;p&gt;The biggest AI stories of this period were not really about feature velocity. They were about dependency.&lt;/p&gt;

&lt;p&gt;Users are beginning to treat these systems as tools they build work and identity around. Once that happens, the decisive questions change. &lt;em&gt;What can it do?&lt;/em&gt; remains important, but it is joined by harder ones:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Can I afford it? Can I trust it? Will it still be available tomorrow? And can I make it mine?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Those were the questions people cared enough to argue about.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>news</category>
      <category>openai</category>
    </item>
    <item>
      <title>The Night AWS Billed the World a Trillion Dollars — and Its Own Alarms Watched It Happen</title>
      <dc:creator>TheRabbitHole</dc:creator>
      <pubDate>Fri, 24 Jul 2026 11:58:40 +0000</pubDate>
      <link>https://dev.to/therabbithole/the-night-aws-billed-the-world-a-trillion-dollars-and-its-own-alarms-watched-it-happen-3dej</link>
      <guid>https://dev.to/therabbithole/the-night-aws-billed-the-world-a-trillion-dollars-and-its-own-alarms-watched-it-happen-3dej</guid>
      <description>&lt;p&gt;Imagine opening your inbox to an AWS Cost Anomaly Detection alert telling you that your side project — the one that costs less than a takeaway coffee each month — has just racked up &lt;strong&gt;$545 million&lt;/strong&gt; in S3 storage charges. Now imagine you weren't alone. On the night of July 16, 2026, thousands of AWS customers around the world got some version of that email. Some saw estimates in the billions. A few saw &lt;em&gt;trillions&lt;/em&gt; — one dashboard reportedly showed $7.1 trillion in month-to-date charges, more than twice Amazon's entire market capitalization.&lt;/p&gt;

&lt;p&gt;The numbers were fake. The panic was not.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually happened
&lt;/h2&gt;

&lt;p&gt;According to the AWS Health Dashboard, the trouble began at roughly 7:46 PM PDT on July 16, when a configuration change in AWS's bill computation system introduced a unit pricing error into the estimated billing pipeline. In plain English: the system that multiplies your usage by a price-per-unit was suddenly multiplying by the wrong unit. Engineers familiar with this class of bug have described how it works — a service meant to charge cents per &lt;em&gt;gigabyte&lt;/em&gt; silently defaults to charging per &lt;em&gt;byte&lt;/em&gt;, and suddenly ordinary usage inflates by a factor of a billion.&lt;/p&gt;

&lt;p&gt;The corrupted numbers flowed downstream into everything that consumes billing estimates: Cost Explorer, AWS Budgets, and — with grim irony — Cost Anomaly Detection, the very service designed to warn customers about runaway spend. Alerts fired en masse. Screenshots flooded social media. One Reddit user posted a bill of $225,579,210,164.83. A developer on X summed up the collective mood: seeing a trillion-dollar figure on your AWS bill is a genuine out-of-body experience, even when some rational part of your brain insists it can't be real.&lt;/p&gt;

&lt;p&gt;Actual invoices were never affected. AWS said so repeatedly, and it was true. But "the number is wrong" is cold comfort at 2 AM when the number has eleven digits.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that should actually worry you
&lt;/h2&gt;

&lt;p&gt;Here's the detail buried in AWS's own timeline that deserves far more attention than the meme-worthy screenshots: &lt;strong&gt;AWS's internal alarms detected the cost anomalies within minutes — and then nothing happened.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;By AWS's own admission, its alarms fired at 7:46 PM PDT on July 16, but they failed to halt the bill generation process and failed to page any engineers. AWS only learned its billing system was broken at 12:19 AM on July 17 — four and a half hours later — &lt;em&gt;because customers told them.&lt;/em&gt; The world's largest cloud provider, the company that sells anomaly detection as a product, discovered its own anomaly through support tickets.&lt;/p&gt;

&lt;p&gt;It gets worse. When an initial rollback of the configuration change failed, AWS paused estimated bill generation entirely — freezing the absurd numbers in place on customer dashboards — and disabled Budgets and Cost Anomaly Detection alerts &lt;em&gt;platform-wide&lt;/em&gt; as a precaution. For the duration of the incident, the two mechanisms AWS tells every customer to rely on for cost safety were switched off globally. Any team that had wired automation to those alerts — Slack escalations, automatic workload shutdowns, spending freezes — was either triggered by phantom data before the pause or flying blind after it.&lt;/p&gt;

&lt;p&gt;The full window of incorrect data lasted roughly 16–24 hours for most customers, with recomputation dragging on longer.&lt;/p&gt;

&lt;h2&gt;
  
  
  A pattern, not a one-off
&lt;/h2&gt;

&lt;p&gt;If this were an isolated stumble, it would be a funny story. It isn't isolated.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;October 2025:&lt;/strong&gt; A race condition in DynamoDB's DNS management system took down us-east-1 for around 15 hours, cascading through 70+ AWS services and knocking out Slack, Snapchat, Ring, Alexa, and huge swaths of the internet. Independent monitoring firm StatusGator crowned us-east-1 the least reliable AWS region of 2025 — 10 outages, nearly 34 hours of cumulative downtime.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;March 2026:&lt;/strong&gt; Another multi-region disruption exposed how many "regional" deployments secretly depend on us-east-1 for authentication and control-plane operations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;May 2026:&lt;/strong&gt; Simultaneous chiller failures in a Northern Virginia data hall triggered a thermal shutdown across EC2 racks — the fourth significant us-east-1 incident in about seven months — fueling speculation that power-hungry AI workloads are straining cooling infrastructure designed for a gentler era.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;July 2026:&lt;/strong&gt; The billing pipeline itself melts down, and the safety systems watch silently.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Trackers noted AWS logged at least one outage every month of 2025 except the final two. Each incident has its own root cause, but the through-line is the same: deeply interconnected systems, hidden single points of failure, and internal guardrails that fail exactly when they're needed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The uncomfortable questions
&lt;/h2&gt;

&lt;p&gt;The billing incident surfaced a fear that goes beyond one bad night. As one Hacker News commenter put it: if AWS can botch billing in a way that produces &lt;em&gt;obviously&lt;/em&gt; impossible numbers, what stops it from botching billing in subtle ways — small overcharges scattered across millions of accounts that nobody would ever notice? Absurd errors announce themselves. Plausible ones just get paid.&lt;/p&gt;

&lt;p&gt;And there's the decade-old grievance the incident reopened: AWS still offers no hard spending cap. You can set alerts. You can set budgets. But you cannot tell AWS "never charge me more than $50, full stop." When the alerts themselves hallucinate — or get switched off platform-wide — customers are left with nothing but trust.&lt;/p&gt;

&lt;p&gt;Some customers decided that trust was exhausted. At least one developer publicly described tearing down all his personal workloads after ten genuinely terrifying minutes of believing a $369 million S3 bill was real, with no plans to return.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you should actually do
&lt;/h2&gt;

&lt;p&gt;A $1.5 trillion estimate should be impossible to display — a basic sanity check would have caught it before it reached a single screen. Since AWS didn't build that check, you should:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Never rely solely on AWS-native cost telemetry.&lt;/strong&gt; Independent FinOps tooling and cross-checked anomaly detection gave teams a second opinion this month when the first opinion went insane.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Alert on actual spend, not just estimates.&lt;/strong&gt; The estimation pipeline and the invoicing pipeline are different systems with different failure modes. July proved the estimates can be fast and catastrophically wrong.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit your automation.&lt;/strong&gt; If a false alert can shut down production workloads, phantom billing data becomes a self-inflicted outage. Add plausibility bounds before any automated response.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assume the safety net can vanish.&lt;/strong&gt; AWS disabled budget alerts globally, mid-incident, without asking. Design as if that can happen again — because it can.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The bigger picture
&lt;/h2&gt;

&lt;p&gt;AWS resolved the issue, apologized, and confirmed nobody would be charged the phantom amounts. What it hasn't done, as of this writing, is publish a full postmortem explaining why its alarms fired into the void, why the rollback failed, or what will change in the billing pipeline's own guardrails.&lt;/p&gt;

&lt;p&gt;For nearly two decades, AWS's brand has rested on a simple promise: &lt;em&gt;we are better at running infrastructure than you are.&lt;/em&gt; Every trillion-dollar phantom bill, every 15-hour regional outage, every alarm that detects a problem and then does nothing chips away at that promise. The cloud isn't collapsing. But the aura of infallibility already has — and in a business built on trust, that might be the more expensive loss.&lt;/p&gt;




&lt;h3&gt;
  
  
  Sources
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://www.infoq.com/news/2026/07/aws-billing-estimates-incident/" rel="noopener noreferrer"&gt;InfoQ — AWS Billing Bug Shows Customers Trillion-Dollar Estimates While Its Own Cost Alarms Fail to Act&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://cybersecuritynews.com/aws-cost-explorer-bug/" rel="noopener noreferrer"&gt;Cybersecurity News — AWS Cost Explorer Bug Shows Trillion-Dollar Billing Estimates&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://gbhackers.com/aws-billing-bug-displays-trillion-dollar-cost-estimates/" rel="noopener noreferrer"&gt;GBHackers — AWS Billing Bug Displays Trillion-Dollar Cost Estimates&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://tech.yahoo.com/computing/articles/amazon-corrects-aws-billing-error-173026819.html" rel="noopener noreferrer"&gt;Yahoo Tech — Amazon Corrects AWS Billing Error Behind Billion-Dollar Invoices&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://statusgator.com/blog/aws-least-reliable-region-in-2025/" rel="noopener noreferrer"&gt;StatusGator — The Least Reliable AWS Region in 2025&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://blog.incidenthub.cloud/definitive-aws-outage-report-2025-reliability" rel="noopener noreferrer"&gt;IncidentHub — The Definitive AWS Outage Report 2025&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.networkworld.com/article/4168878/aws-hit-by-us-east-1-outage-after-data-center-thermal-event.html" rel="noopener noreferrer"&gt;Network World — AWS hit by US-East-1 outage after data center thermal event&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
    </item>
    <item>
      <title>I Asked My AI "Who Are You" — It Cost 42,000 Tokens to Answer</title>
      <dc:creator>TheRabbitHole</dc:creator>
      <pubDate>Fri, 24 Jul 2026 09:43:22 +0000</pubDate>
      <link>https://dev.to/therabbithole/-i-asked-my-ai-who-are-you-it-cost-42000-tokens-to-answer-44jp</link>
      <guid>https://dev.to/therabbithole/-i-asked-my-ai-who-are-you-it-cost-42000-tokens-to-answer-44jp</guid>
      <description>&lt;p&gt;I typed a 19-character question into Claude Code:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;who are you?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then I opened the network logs of what my machine actually sent over the wire. I recommend you never do this. It ruins something in you.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;p&gt;My question was &lt;strong&gt;19 bytes&lt;/strong&gt;. The request that carried it was &lt;strong&gt;118,693 bytes&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That's a 6,200x amplification before the model reads a single word I wrote. My actual message was &lt;strong&gt;0.016%&lt;/strong&gt; of the payload. If this were a letter, I'd be mailing a phone book with a Post-it note inside.&lt;/p&gt;

&lt;p&gt;Here's the pie, byte for byte:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What&lt;/th&gt;
&lt;th&gt;Bytes&lt;/th&gt;
&lt;th&gt;Share&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;38 tool JSON schemas&lt;/td&gt;
&lt;td&gt;78,935&lt;/td&gt;
&lt;td&gt;66.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;System prompt (behavior manual, safety rules, my git status)&lt;/td&gt;
&lt;td&gt;22,178&lt;/td&gt;
&lt;td&gt;18.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A list of 60 &lt;em&gt;more&lt;/em&gt; tools it could load, plus 25 skill descriptions&lt;/td&gt;
&lt;td&gt;16,844&lt;/td&gt;
&lt;td&gt;14.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;My memory files, my email address, today's date&lt;/td&gt;
&lt;td&gt;1,126&lt;/td&gt;
&lt;td&gt;1.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;My question&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;19&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.016%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;And the tool schemas aren't even schemas, mostly. Of those 78,935 bytes, only 29,164 are machine-readable JSON schema. The other &lt;strong&gt;44,392 bytes are English prose&lt;/strong&gt; — instruction manuals for the model, shipped on every single request. One tool alone, &lt;code&gt;Workflow&lt;/code&gt;, carries a 19,012-byte description. That's a small essay. It rides along on every message I send, forever, whether I orchestrate a multi-agent workflow or just say "hi."&lt;/p&gt;

&lt;p&gt;Sixteen of the 38 tools are for driving a browser preview I never opened. There's an iOS Simulator controller in there. I was asking a chatbot its name.&lt;/p&gt;

&lt;h2&gt;
  
  
  The full manifest, since you asked
&lt;/h2&gt;

&lt;p&gt;Every tool schema in that request, largest first. "Prose" is the English description; "schema" is the actual machine-readable part. Approximate tokens = bytes ÷ 4.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Bytes&lt;/th&gt;
&lt;th&gt;Prose&lt;/th&gt;
&lt;th&gt;Schema&lt;/th&gt;
&lt;th&gt;~Tokens&lt;/th&gt;
&lt;th&gt;Did I need it?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Workflow (multi-agent orchestration)&lt;/td&gt;
&lt;td&gt;21,609&lt;/td&gt;
&lt;td&gt;19,012&lt;/td&gt;
&lt;td&gt;1,833&lt;/td&gt;
&lt;td&gt;~5,400&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Artifact (publish web pages)&lt;/td&gt;
&lt;td&gt;10,679&lt;/td&gt;
&lt;td&gt;6,358&lt;/td&gt;
&lt;td&gt;4,002&lt;/td&gt;
&lt;td&gt;~2,670&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AskUserQuestion&lt;/td&gt;
&lt;td&gt;4,371&lt;/td&gt;
&lt;td&gt;1,097&lt;/td&gt;
&lt;td&gt;3,149&lt;/td&gt;
&lt;td&gt;~1,090&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;iOS Simulator control&lt;/td&gt;
&lt;td&gt;3,976&lt;/td&gt;
&lt;td&gt;1,307&lt;/td&gt;
&lt;td&gt;2,518&lt;/td&gt;
&lt;td&gt;~990&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ScheduleWakeup (self-pacing loops)&lt;/td&gt;
&lt;td&gt;3,942&lt;/td&gt;
&lt;td&gt;2,685&lt;/td&gt;
&lt;td&gt;1,068&lt;/td&gt;
&lt;td&gt;~985&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;visualize show_widget&lt;/td&gt;
&lt;td&gt;3,038&lt;/td&gt;
&lt;td&gt;648&lt;/td&gt;
&lt;td&gt;2,259&lt;/td&gt;
&lt;td&gt;~760&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent (spawn subagents)&lt;/td&gt;
&lt;td&gt;3,037&lt;/td&gt;
&lt;td&gt;1,574&lt;/td&gt;
&lt;td&gt;1,342&lt;/td&gt;
&lt;td&gt;~760&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bash&lt;/td&gt;
&lt;td&gt;2,860&lt;/td&gt;
&lt;td&gt;1,290&lt;/td&gt;
&lt;td&gt;1,456&lt;/td&gt;
&lt;td&gt;~715&lt;/td&gt;
&lt;td&gt;No (but fair)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ReportFindings (code review UI)&lt;/td&gt;
&lt;td&gt;2,314&lt;/td&gt;
&lt;td&gt;574&lt;/td&gt;
&lt;td&gt;1,646&lt;/td&gt;
&lt;td&gt;~580&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Browser computer (mouse/keyboard)&lt;/td&gt;
&lt;td&gt;2,141&lt;/td&gt;
&lt;td&gt;193&lt;/td&gt;
&lt;td&gt;1,837&lt;/td&gt;
&lt;td&gt;~535&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Skill&lt;/td&gt;
&lt;td&gt;1,699&lt;/td&gt;
&lt;td&gt;1,244&lt;/td&gt;
&lt;td&gt;345&lt;/td&gt;
&lt;td&gt;~425&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Read&lt;/td&gt;
&lt;td&gt;1,670&lt;/td&gt;
&lt;td&gt;790&lt;/td&gt;
&lt;td&gt;776&lt;/td&gt;
&lt;td&gt;~420&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Browser preview_start&lt;/td&gt;
&lt;td&gt;1,656&lt;/td&gt;
&lt;td&gt;1,328&lt;/td&gt;
&lt;td&gt;143&lt;/td&gt;
&lt;td&gt;~415&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;spawn_task (background task chips)&lt;/td&gt;
&lt;td&gt;1,553&lt;/td&gt;
&lt;td&gt;664&lt;/td&gt;
&lt;td&gt;765&lt;/td&gt;
&lt;td&gt;~390&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ToolSearch (loads MORE tools)&lt;/td&gt;
&lt;td&gt;1,522&lt;/td&gt;
&lt;td&gt;953&lt;/td&gt;
&lt;td&gt;427&lt;/td&gt;
&lt;td&gt;~380&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;mark_chapter (session table of contents)&lt;/td&gt;
&lt;td&gt;1,092&lt;/td&gt;
&lt;td&gt;699&lt;/td&gt;
&lt;td&gt;256&lt;/td&gt;
&lt;td&gt;~275&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Edit&lt;/td&gt;
&lt;td&gt;1,037&lt;/td&gt;
&lt;td&gt;360&lt;/td&gt;
&lt;td&gt;584&lt;/td&gt;
&lt;td&gt;~260&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;dismiss_task&lt;/td&gt;
&lt;td&gt;918&lt;/td&gt;
&lt;td&gt;560&lt;/td&gt;
&lt;td&gt;234&lt;/td&gt;
&lt;td&gt;~230&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;+ 20 more (browser tabs, navigation, console readers, form filler, widget reader, Write…)&lt;/td&gt;
&lt;td&gt;~9,800&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;~2,450&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;78,935&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;44,392&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;29,164&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~19,700&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0 of 38 used&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Read that last column again. The model answered my question using &lt;strong&gt;zero&lt;/strong&gt; of the 38 tools it was handed. Then notice &lt;code&gt;ToolSearch&lt;/code&gt;: a tool whose only purpose is to fetch &lt;em&gt;even more tools&lt;/em&gt; — there are ~60 additional ones (WebSearch, WebFetch, cron schedulers, task managers, notebook editors, a whole second browser that drives my real Chrome, password-manager integration) waiting behind it, listed by name in that 16.8KB deferred-tools block. The request contains a tool for requesting tools. We have achieved tool-calling recursion.&lt;/p&gt;

&lt;p&gt;The prose-to-schema ratio is the tell. &lt;code&gt;Workflow&lt;/code&gt;'s description is 19KB of instruction manual — pipeline semantics, barrier vs. no-barrier philosophy, seven design patterns with code examples, a "smell test" for when you're using it wrong. It's genuinely well-written documentation. It is also transmitted to a datacenter every time I type "ok".&lt;/p&gt;

&lt;h2&gt;
  
  
  What you can actually turn off (and what you can't)
&lt;/h2&gt;

&lt;p&gt;I audited my own machine after this. Here's the honest control surface, verified against my actual config files — with guesses labeled as guesses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Things I verified you CAN control:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your own hooks.&lt;/strong&gt; Mine injects a paragraph of instructions into &lt;em&gt;every&lt;/em&gt; prompt I send (my doing, &lt;code&gt;~/.claude/settings.json&lt;/code&gt;). Delete the hook block, it's gone.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory files.&lt;/strong&gt; Plain markdown in &lt;code&gt;~/.claude/projects/&amp;lt;project&amp;gt;/memory/&lt;/code&gt;. Delete them, they stop being sent. (The app injecting my email address alongside them — that part has no knob.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plugins.&lt;/strong&gt; My skill list was fattened by exactly one thing I installed: the official &lt;code&gt;anthropic-skills&lt;/code&gt; plugin. Uninstalling it removes ~25 skill descriptions from every request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Telemetry env vars.&lt;/strong&gt; &lt;code&gt;DISABLE_TELEMETRY=1&lt;/code&gt;, &lt;code&gt;DISABLE_ERROR_REPORTING=1&lt;/code&gt;, &lt;code&gt;CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1&lt;/code&gt; are documented switches for the ~600KB &lt;code&gt;tengu_*&lt;/code&gt; firehose. (Documented — I haven't packet-captured a session with them on yet. Trust, but HAR-verify.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The entrypoint.&lt;/strong&gt; The desktop app injects ~20 MCP tool schemas (browser preview, iOS simulator, visualization widgets, session chips) that plain terminal &lt;code&gt;claude&lt;/code&gt; doesn't load. Using the terminal is the single biggest schema diet available to a normal user.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Things you CANNOT control, period:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The core toolset.&lt;/strong&gt; Workflow's 21.6KB, Artifact's 10.7KB, all of it — baked into the binary, gated by &lt;em&gt;server-side&lt;/em&gt; feature flags (&lt;code&gt;tengu_workflows_enabled: true&lt;/code&gt;, &lt;code&gt;tengu_report_findings_tool: true&lt;/code&gt;, &lt;code&gt;tengu_chrome_auto_enable: true&lt;/code&gt;). There is a &lt;code&gt;--disallowedTools&lt;/code&gt; launch flag on the CLI, but the desktop app exposes no such thing in its UI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The 20.7KB system prompt.&lt;/strong&gt; Not a config file anywhere. Ships with the binary, changes when Anthropic says so.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The ~370 feature flags.&lt;/strong&gt; They live server-side; your local copy (&lt;code&gt;.claude.json&lt;/code&gt;) is a cache that re-syncs on their schedule. I found flags in mine for experiments I've never seen, promo banners, A/B copy tests, per-model "velvet hammer" variants (twenty-two of those, whatever they are), and two remote kill switches — &lt;code&gt;tengu-off-switch&lt;/code&gt; and &lt;code&gt;tengu-fable-off-switch&lt;/code&gt; — that Anthropic can flip and I cannot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The title-generation side call.&lt;/strong&gt; No setting. Every new session, a second model gets paid to name your conversation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The 40-call self-measurement burst.&lt;/strong&gt; No setting. The context meter wants what it wants.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your identity riding along.&lt;/strong&gt; Account UUID, hashed device ID, session ID go in the API &lt;code&gt;metadata&lt;/code&gt; of every model call. OS, arch, node version, shell, package managers go in the telemetry. None of it optional.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The asymmetry is the story. The user's levers — a hook, some markdown files, one plugin, an env var — govern maybe 3% of the payload. The other 97% is decided by people you've never met, toggled by flags with names like &lt;code&gt;tengu_amber_moleskin&lt;/code&gt;, shipped to you on every keystroke, and cached with a straight face.&lt;/p&gt;

&lt;h2&gt;
  
  
  The answer took 8.7 seconds. Here's where they went.
&lt;/h2&gt;

&lt;p&gt;The response to my question was &lt;strong&gt;127 tokens&lt;/strong&gt; — 47 of them the model privately thinking "user's being blunt, keep it brief," and 80 tokens of actual answer.&lt;/p&gt;

&lt;p&gt;Generating those 127 tokens takes well under a second. The other ~8 seconds? The model had to ingest &lt;strong&gt;42,270 tokens of context&lt;/strong&gt; first — and 13,445 of them weren't in the prompt cache yet, because this was the session's first message. I paid an eight-second cold-start tax so the model could re-read the operating manual for a browser I never launched.&lt;/p&gt;

&lt;p&gt;The economics in one line: &lt;strong&gt;42,270 tokens in, 127 tokens out.&lt;/strong&gt; A 333:1 read-to-write ratio. To say hello.&lt;/p&gt;

&lt;h2&gt;
  
  
  It gets weirder
&lt;/h2&gt;

&lt;p&gt;While answering me, the app quietly made a &lt;strong&gt;second, parallel model call&lt;/strong&gt; — to a different, smaller model — whose entire job was to invent a title for the sidebar. Input: my rude question. Output: &lt;code&gt;"Clarify Claude agent identity"&lt;/code&gt;. Five hundred twenty-one tokens of prompt to produce four polite words I never asked for.&lt;/p&gt;

&lt;p&gt;After the answer landed, the client fired &lt;strong&gt;40 more API calls in one burst&lt;/strong&gt; — all to a token-counting endpoint. Why 40? It was measuring itself: one call per slice of its own system prompt, plus one call per tool sending the string &lt;code&gt;"foo"&lt;/code&gt; next to a single schema, so it could subtract the baseline and learn each tool's exact token weight. The app literally runs a scientific experiment on its own bloat, every session, to render a little context meter.&lt;/p&gt;

&lt;p&gt;And the telemetry. Oh, the telemetry. In the same 15-second window: &lt;strong&gt;~600KB across 184+ events&lt;/strong&gt;, every one namespaced &lt;code&gt;tengu_&lt;/code&gt; — Anthropic's internal codename for Claude Code. &lt;code&gt;tengu_feature_ok&lt;/code&gt; fired 72 times, once per remote feature-flag check. &lt;code&gt;tengu_skill_loaded&lt;/code&gt; fired 48 times. The telemetry shipped to log my question &lt;em&gt;outweighed the question&lt;/em&gt; by a factor of thirty thousand. (To be fair: I grepped every event — my actual words appear in none of them.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Who configured this? Nobody. That's the point.
&lt;/h2&gt;

&lt;p&gt;I went looking for the config file where I'd opted into all this. It doesn't exist. My settings contain a theme and one hook. My MCP server list is empty in every project. Every fat tool, every schema, every deferred-tool stub is switched on by &lt;strong&gt;~370 server-side feature flags&lt;/strong&gt; with names like &lt;code&gt;tengu_velvet_hammer_opus_4_8&lt;/code&gt; and &lt;code&gt;tengu_amber_moleskin&lt;/code&gt;, cached to my disk and re-synced whenever the mothership pleases. There's even a complete 5KB onboarding-guide prompt template embedded &lt;em&gt;inside the flag cache&lt;/em&gt;, sitting on my SSD in case a feature I've never used ever wants it.&lt;/p&gt;

&lt;p&gt;There are two kill switches in that file: &lt;code&gt;tengu-off-switch&lt;/code&gt; and &lt;code&gt;tengu-fable-off-switch&lt;/code&gt;. Neither is mine to flip.&lt;/p&gt;

&lt;h2&gt;
  
  
  The uncomfortable defense
&lt;/h2&gt;

&lt;p&gt;Here's the part that annoys me most: it's not even stupid. The 100KB of scaffolding is prompt-cached — I pay full prefill once per session, then it's cheap cache reads. The deferred-tool list I mocked exists precisely to &lt;em&gt;avoid&lt;/em&gt; sending 60 more schemas. The 8.7 seconds is a one-time tax; turn two is fast. Every individual decision is defensible. The system is locally rational and globally deranged — a machine that ships a library with every postcard and has optimized the shipping.&lt;/p&gt;

&lt;p&gt;But step back. A human answering "who are you" needs zero context refresh, no tool manifest, no telemetry burst, no second brain generating a title. We've built agents that must be handed their entire identity, capabilities, safety rules, and worldview from scratch on every request, because the underlying model remembers nothing. Statelessness is the original sin; everything else — the caching, the deferral, the 40-call self-measurement — is increasingly elaborate penance.&lt;/p&gt;

&lt;p&gt;19 bytes in. 118,693 bytes of context, 600KB of telemetry, 43 API calls, two models, and 8.7 seconds later: "I'm Claude Code."&lt;/p&gt;

&lt;p&gt;I know. I read the packet capture.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>What Actually Matters When You're Hunting a Generative AI Job</title>
      <dc:creator>TheRabbitHole</dc:creator>
      <pubDate>Fri, 10 Jul 2026 14:08:28 +0000</pubDate>
      <link>https://dev.to/therabbithole/what-actually-matters-when-youre-hunting-a-generative-ai-job-pcc</link>
      <guid>https://dev.to/therabbithole/what-actually-matters-when-youre-hunting-a-generative-ai-job-pcc</guid>
      <description>&lt;p&gt;Everybody keeps saying the Generative AI job market is on fire. Fine — but almost nobody tells you what those jobs &lt;em&gt;actually&lt;/em&gt; ask for once you get past the buzzwords. So I stopped guessing and pulled &lt;strong&gt;95 live "Generative AI" postings from Google Jobs (US)&lt;/strong&gt;, then counted what really shows up in the titles and the full descriptions. The interesting part wasn't the individual skills — it was which ones keep showing up &lt;em&gt;together&lt;/em&gt;. That's the bit that should change how you prep.&lt;/p&gt;

&lt;p&gt;Here's what I found, receipts included.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;A "Generative AI" job today is mostly an &lt;strong&gt;applied engineering&lt;/strong&gt; job: building with existing models, not training new foundation models from scratch.&lt;/li&gt;
&lt;li&gt;Everything is converging on one thing: &lt;strong&gt;agentic RAG systems&lt;/strong&gt;. RAG and agents barely show up apart anymore — 78% of the "agent" postings also mention RAG.&lt;/li&gt;
&lt;li&gt;The &lt;strong&gt;core stack — LLM + Python + RAG — appears together in ~36% of postings (34 of 95)&lt;/strong&gt;, the highest-leverage combo in the sample. Layer on agents and an orchestration framework and you cover most of the &lt;em&gt;frequently recurring&lt;/em&gt; combinations.&lt;/li&gt;
&lt;li&gt;The money is real: among the &lt;strong&gt;23 postings that disclosed pay&lt;/strong&gt;, the median range-midpoint was about &lt;strong&gt;$187k&lt;/strong&gt;, with individual midpoints spanning &lt;strong&gt;$121k–$274k&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Reality check nobody hands you: only &lt;strong&gt;13%&lt;/strong&gt; of these roles are at Big Tech. Nearly half the market is a long tail of companies you've genuinely never heard of.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;One caveat on scope:&lt;/strong&gt; this is &lt;strong&gt;95 US postings&lt;/strong&gt; from Google Jobs, captured mid-2026. Treat it as a &lt;em&gt;trend read&lt;/em&gt;, not a census. The US is the biggest, most mature GenAI labor market, so it's a decent leading indicator — but the numbers will look different in the EU, India, and elsewhere. Directional, not gospel.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And one on method:&lt;/strong&gt; every percentage below measures what employers &lt;strong&gt;foreground in a description&lt;/strong&gt;, not a verified hard requirement. A term can appear as "required," "a plus," or even "the legacy thing we're replacing." Read these as &lt;em&gt;what the market talks about&lt;/em&gt;, not a checklist every employer enforces. (Full methodology at the end.)&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The skill tiers: what to learn, what to skip
&lt;/h2&gt;

&lt;p&gt;I split every skill by how often it shows up across the 95 postings. Read this as signal strength, not a set of hard gates:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;Skills&lt;/th&gt;
&lt;th&gt;What it means for you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Core signals (&amp;gt;50%)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;LLMs (75%), Python (61%), RAG (61%), Prompt engineering (54%), Agents (54%)&lt;/td&gt;
&lt;td&gt;Appear in the majority of postings — the safest foundational bets&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Common differentiators (15–40%)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Fine-tuning (38%), LangChain (34%), MLOps (22%), Vector DBs (22%), Multimodal (20%)&lt;/td&gt;
&lt;td&gt;Frequently useful, less universal — these break ties in your favor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Narrower signals (&amp;lt;15%)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Role-specific tooling, niche domains&lt;/td&gt;
&lt;td&gt;Relevant to specific segments, not the broad market&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One thing sits &lt;em&gt;outside&lt;/em&gt; the skill tiers because it isn't a skill: a &lt;strong&gt;US security clearance&lt;/strong&gt; appears in ~13% of postings. It's an eligibility/access requirement, not something you learn — but if you have one, it's a genuinely low-competition lane (more below).&lt;/p&gt;

&lt;p&gt;If you're short on time, get the core-signal five solid &lt;em&gt;first&lt;/em&gt;. They're common enough to be the highest-expected-value things to know — a deep grasp of fine-tuning won't help much if you can't stand up a RAG pipeline in Python.&lt;/p&gt;

&lt;h2&gt;
  
  
  The thing the raw keyword counts hide: RAG and agents are now one skill
&lt;/h2&gt;

&lt;p&gt;This is the finding I most want you to sit with. Look at which skills land in the &lt;em&gt;same&lt;/em&gt; posting:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Of the &lt;strong&gt;51&lt;/strong&gt; postings that mention "agent," &lt;strong&gt;40&lt;/strong&gt; also mention &lt;strong&gt;RAG&lt;/strong&gt; (78%).&lt;/li&gt;
&lt;li&gt;Of the &lt;strong&gt;58&lt;/strong&gt; that mention &lt;strong&gt;RAG&lt;/strong&gt;, that same &lt;strong&gt;40&lt;/strong&gt; also mention agents (69%).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;92%&lt;/strong&gt; of agent postings also mention &lt;strong&gt;LLMs&lt;/strong&gt;; &lt;strong&gt;51%&lt;/strong&gt; specifically name &lt;strong&gt;LangChain&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;66%&lt;/strong&gt; of RAG postings also mention &lt;strong&gt;prompt engineering&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So there's a hard core of ~&lt;strong&gt;40 postings — about 42% of the whole sample&lt;/strong&gt; — that foreground RAG &lt;em&gt;and&lt;/em&gt; agents together. The market isn't describing "a RAG person" or "an agent person." It's describing someone who ships &lt;strong&gt;agentic RAG&lt;/strong&gt; — retrieval-grounded systems that then go &lt;em&gt;do things&lt;/em&gt; through tool calls and multi-step planning. Build one project that does both and you're suddenly speaking the exact language most of these listings are written in.&lt;/p&gt;

&lt;h2&gt;
  
  
  The stack you should actually build
&lt;/h2&gt;

&lt;p&gt;Frameworks matter, so let's name names instead of waving our hands at "AI tools":&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Orchestration:&lt;/strong&gt; &lt;strong&gt;LangChain (34%)&lt;/strong&gt; is the default. &lt;strong&gt;LlamaIndex (11%)&lt;/strong&gt; is the RAG-specialist pick. 34% name at least one — so learn LangChain first, and pick up LlamaIndex too if the role is retrieval-heavy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model providers:&lt;/strong&gt; &lt;strong&gt;OpenAI (29%)&lt;/strong&gt; leads; &lt;strong&gt;Anthropic/Claude (~12%)&lt;/strong&gt; is the clear, growing #2. Honestly, employers care less about &lt;em&gt;which&lt;/em&gt; SDK and more that you've dealt with the ugly parts — token limits, streaming, evals, latency, cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieval layer:&lt;/strong&gt; vector DBs + embeddings show up in ~22%. Know how chunking, embedding models, and a vector store (pgvector, Pinecone, Weaviate, take your pick) actually fit together.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ship it:&lt;/strong&gt; MLOps (22%), Docker (22%), Kubernetes (21%). A model sitting in a notebook doesn't get you hired. A deployed endpoint does.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Languages:&lt;/strong&gt; Python (61%) isn't optional. Java (14%) and TypeScript/React (~6–11%) turn up in the more product-facing roles.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The one portfolio project that covers most of this:&lt;/strong&gt; a Python service that does RAG over a real corpus (chunk → embed → vector store → retrieve), wraps it in an agent with tool use via LangChain, calls OpenAI or Claude, has some basic evals, and ships in a Docker container. That single build hits the core-signal five &lt;em&gt;plus&lt;/em&gt; three differentiators.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pick a cloud — but it's not winner-take-all
&lt;/h2&gt;

&lt;p&gt;Cloud isn't optional, and the order is clear: &lt;strong&gt;AWS (40%) &amp;gt; Azure (35%) &amp;gt; GCP (25%)&lt;/strong&gt;, with Vertex AI showing up in ~9%. Here's the twist though: &lt;strong&gt;29% of postings mention 2+ clouds.&lt;/strong&gt; Multi-cloud fluency is a differentiator all by itself. Starting from zero? AWS has the widest coverage here, so it's the safe default — but for Microsoft-heavy enterprise and consulting shops, Azure is often the smarter first pick. Then get comfortable in a second.&lt;/p&gt;

&lt;h2&gt;
  
  
  Titles and seniority: there's room in the middle
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Engineer&lt;/strong&gt; (34) is the default title, then &lt;strong&gt;Scientist/Research&lt;/strong&gt; (15), &lt;strong&gt;Developer&lt;/strong&gt; (10), &lt;strong&gt;Architect&lt;/strong&gt; (8), and &lt;strong&gt;Manager/Lead&lt;/strong&gt; (8).&lt;/li&gt;
&lt;li&gt;Only ~18% are explicitly "Senior." Most postings don't even bother specifying a level.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Translation: you don't need "Staff" on your résumé to get a foot in the door. Mid-level GenAI hiring is wide open right now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the jobs actually are: clustering the employers
&lt;/h2&gt;

&lt;p&gt;I bucketed all 95 companies into rough sectors by company identity. The buckets are approximate and not perfectly exclusive — a defense consultancy could arguably sit in two — but the shape is the expectations reset almost nobody gives you up front:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Company cluster&lt;/th&gt;
&lt;th&gt;Share&lt;/th&gt;
&lt;th&gt;Examples&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Other mid-market, specialist &amp;amp; less-recognized employers&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~47%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Mid-size firms, boutiques, contractors, subsidiaries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Big Tech / enterprise product&lt;/td&gt;
&lt;td&gt;~13%&lt;/td&gt;
&lt;td&gt;Adobe, Oracle, NVIDIA, AT&amp;amp;T, DoorDash&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Consulting &amp;amp; system integrators&lt;/td&gt;
&lt;td&gt;~9%&lt;/td&gt;
&lt;td&gt;Deloitte, Booz Allen, c1advantage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AI-native companies / startups&lt;/td&gt;
&lt;td&gt;~9%&lt;/td&gt;
&lt;td&gt;Kendia.AI, Innodata, EnthuZiastic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Finance &amp;amp; banking&lt;/td&gt;
&lt;td&gt;~5%&lt;/td&gt;
&lt;td&gt;JPMorgan, capital-markets firms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Staffing / recruiting agencies&lt;/td&gt;
&lt;td&gt;~5%&lt;/td&gt;
&lt;td&gt;US Tech Solutions, SVAM, Saransh&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Semiconductors / hardware&lt;/td&gt;
&lt;td&gt;~4%&lt;/td&gt;
&lt;td&gt;Sandisk, Applied Materials, Infinite Electronics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Academia&lt;/td&gt;
&lt;td&gt;~4%&lt;/td&gt;
&lt;td&gt;Harvard and university labs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Defense primes&lt;/td&gt;
&lt;td&gt;~2%&lt;/td&gt;
&lt;td&gt;Peraton (plus the clearance roles below)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things jump out here:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The long tail &lt;em&gt;is&lt;/em&gt; the market.&lt;/strong&gt; Around 47% of postings are at companies that aren't consumer-recognizable brands — some are large contractors or subsidiaries, just not names you'd rattle off. If you only apply to FAANG and the hot AI startups, you're fighting over roughly &lt;strong&gt;22%&lt;/strong&gt; of the openings — against everyone else doing the exact same thing. The real volume, and less competition, lives in the mid-market, the consultancies, and the contractors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Clearance is a parallel market.&lt;/strong&gt; About &lt;strong&gt;13%&lt;/strong&gt; of roles want a US security clearance (that cuts across defense primes, consultancies, and gov contractors). Got one? It's your fastest, least-crowded way in. Don't have one? Filter those out and stop burning applications on them.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If your mental image of a "GenAI job" is a FAANG research lab, recalibrate. Write your résumé for &lt;strong&gt;delivery and business impact&lt;/strong&gt; — that's what the consultancies and mid-market shops are actually buying — not just model benchmarks.&lt;/p&gt;

&lt;p&gt;One more quiet signal from the listings: &lt;strong&gt;remote-first roles were uncommon in this sample&lt;/strong&gt; — only ~10% prominently advertised remote. That doesn't automatically make the rest on-site (plenty are hybrid or just don't say), but if you're remote-only, you're fishing in a smaller pond here. Plan for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your actual to-do list
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Build one agentic RAG app&lt;/strong&gt;, end to end (Python + vector DB + LangChain + OpenAI/Claude + Docker). That single project matches the core-signal five and several differentiators.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Add evals and cost/latency handling&lt;/strong&gt; — the boring 20% that separates a demo from a hire.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pick a cloud that matches your target sector.&lt;/strong&gt; AWS has the broadest coverage here (40%), but Azure is often the better first bet for Microsoft-heavy enterprise and consulting roles. Then get literate in a second.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Apply to mid-level roles broadly.&lt;/strong&gt; Don't self-filter for "Senior."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide on the clearance lane.&lt;/strong&gt; If you've got one, lead with it — it's your fastest, least-competitive path in.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The GenAI market rewards people who've actually &lt;em&gt;built and shipped&lt;/em&gt;, not people who've &lt;em&gt;read about&lt;/em&gt; transformers. Pick one end-to-end agentic-RAG project, make it real, and let these 95 job descriptions be your syllabus.&lt;/p&gt;




&lt;h2&gt;
  
  
  How I analyzed the postings
&lt;/h2&gt;

&lt;p&gt;Because this whole post is built on original data, here's exactly what I did so you can weigh it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Source &amp;amp; date:&lt;/strong&gt; Scraped from Google Jobs on &lt;strong&gt;2026-07-10&lt;/strong&gt; via a scraper Actor. Query &lt;strong&gt;"Generative AI"&lt;/strong&gt;, location &lt;strong&gt;United States&lt;/strong&gt;, country &lt;strong&gt;US&lt;/strong&gt; — a fixed query, &lt;em&gt;not&lt;/em&gt; personalized to my own location or search history.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sample:&lt;/strong&gt; &lt;strong&gt;95 postings&lt;/strong&gt; returned. A mix of full-time, contract, and part-time. I did &lt;strong&gt;not&lt;/strong&gt; de-duplicate repeat employers, so a few companies appear more than once (e.g. Innodata ×4, Deloitte ×3) — that reflects real posting volume but slightly weights active hirers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keyword prevalence:&lt;/strong&gt; case-insensitive &lt;strong&gt;substring match&lt;/strong&gt; across each posting's title + full description. Variants group naturally — "agent" catches &lt;em&gt;agents/agentic/AI agents&lt;/em&gt;; "fine-tun" catches &lt;em&gt;fine-tune/fine-tuning&lt;/em&gt;. Each percentage is the share of the 95 postings whose text contains the term.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mention ≠ requirement.&lt;/strong&gt; This measures what employers &lt;em&gt;foreground in the text&lt;/em&gt;, not verified hard requirements. A term can appear as required, "a plus," or context. Treat the numbers as what the market talks about.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Co-occurrence&lt;/strong&gt; counts are raw intersections (e.g. 40 of 95 postings contain both "agent" and "RAG").&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Salary:&lt;/strong&gt; only &lt;strong&gt;23 of 95&lt;/strong&gt; postings disclosed a range. I took each range's &lt;strong&gt;midpoint&lt;/strong&gt;; the reported figures are the &lt;strong&gt;median&lt;/strong&gt; of those midpoints ($187.5k) and their &lt;strong&gt;min/max&lt;/strong&gt; ($121.2k–$274.1k). This subset may be biased toward employers, states, or role types more likely to publish pay.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Employer sectors:&lt;/strong&gt; assigned by company identity; buckets are approximate and non-exclusive. The ~47% "other" bucket = companies that didn't match a named sector — not a claim that they're small.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Known limits:&lt;/strong&gt; single query, single day, US-only, one job board. Good for an &lt;strong&gt;exploratory trend read&lt;/strong&gt;; not a census of the GenAI labor market.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>career</category>
      <category>llm</category>
      <category>rag</category>
    </item>
    <item>
      <title>Project Glasswing: The Death Verdict for Open Source?</title>
      <dc:creator>TheRabbitHole</dc:creator>
      <pubDate>Fri, 10 Apr 2026 09:06:08 +0000</pubDate>
      <link>https://dev.to/therabbithole/project-glasswing-and-the-mythos-moment-a-critical-examination-of-ais-cybersecurity-crossroads-129d</link>
      <guid>https://dev.to/therabbithole/project-glasswing-and-the-mythos-moment-a-critical-examination-of-ais-cybersecurity-crossroads-129d</guid>
      <description>&lt;p&gt;On April 7, 2026, Anthropic announced Project Glasswing—a defensive cybersecurity initiative built around Claude Mythos Preview, a frontier AI model so capable at finding and exploiting vulnerabilities that Anthropic deems it too dangerous for general public release. Backed by $100 million in usage credits and a "coalition of the willing" including Amazon, Apple, Google, Microsoft, Nvidia, the Linux Foundation, CrowdStrike, Palo Alto Networks, and more, Project Glasswing aims to give defenders a head start before similar capabilities proliferate to adversarial actors.&lt;/p&gt;

&lt;p&gt;The announcement arrived during a remarkable week for Anthropic: the company disclosed $30 billion in annualized revenue (tripling in months), sealed a multi-gigawatt compute deal with Google and Broadcom, and faces potential IPO considerations. This timing raises immediate questions about whether Glasswing represents a watershed moment for cybersecurity, a strategic business move, or both.&lt;/p&gt;

&lt;p&gt;What follows is a deep investigation drawing on Anthropic's own documentation, independent press analysis, technical community response, and security expert perspectives to evaluate Project Glasswing—the claims, the risks, the business strategy, and what it means for the future of digital security.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Capabilities: Something Remarkable, or Marketing Hyperbole?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What Anthropic Claims
&lt;/h3&gt;

&lt;p&gt;According to Anthropic's comprehensive technical evaluation, Claude Mythos Preview demonstrates:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Autonomous discovery of &lt;strong&gt;thousands of zero-day vulnerabilities&lt;/strong&gt; in every major operating system and web browser&lt;/li&gt;
&lt;li&gt;Ability to develop &lt;strong&gt;working exploits without human intervention&lt;/strong&gt;—in one case chaining together four distinct vulnerabilities to escape browser sandboxes&lt;/li&gt;
&lt;li&gt;Spectacular benchmark results: &lt;strong&gt;83.1%&lt;/strong&gt; on CyberGym versus 66.6% for Claude Opus 4.6, and &lt;strong&gt;93.9%&lt;/strong&gt; on SWE-bench Verified&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Particularly striking are specific examples:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A &lt;strong&gt;27-year-old vulnerability&lt;/strong&gt; in OpenBSD—a security-focused OS—that allowed remote crash by mere connection&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;16-year-old bug&lt;/strong&gt; in FFmpeg's H.264 codec, surviving five million automated fuzzing attempts&lt;/li&gt;
&lt;li&gt;Autonomous &lt;strong&gt;local privilege escalation exploits&lt;/strong&gt; on Linux by chaining multiple vulnerabilities&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  External Verification
&lt;/h3&gt;

&lt;p&gt;FFmpeg maintainers have confirmed patches were submitted noting they "appear to be written by humans." Greg Kroah-Hartman, the Linux stable maintainer, has publicly stated: "Months ago, we were getting 'AI slop'... Something happened a month ago, and the world switched. Now we have real reports." Security teams across major open source projects report the same shift.&lt;/p&gt;

&lt;p&gt;Forbes analyst Paulo Carvão notes that the evidence is "difficult to dismiss" given that Mythos can "chain together vulnerabilities that individually appear benign but collectively yield complete system compromise."&lt;/p&gt;

&lt;h3&gt;
  
  
  The Skeptical Community Response
&lt;/h3&gt;

&lt;p&gt;On Hacker News, responses range from excitement about genuine advancement to bitter skepticism about relentless "doomer" marketing. One security professional noted they've already had success using existing models: "I've had these successes without scaffolding or really anything past Claude CLI and a small prompt as well? So like I'm in a weird place where this was already happening and Mythos is being sold like it wasn't good before?"&lt;/p&gt;

&lt;p&gt;Others point out that we've heard dramatic breakthrough claims before. Anthropic's own CEO previously claimed 90% of code would be written by LLMs within 3-6 months—a timeline clearly not met. There's fatigue with each iteration being framed as world-endingly powerful.&lt;/p&gt;

&lt;h3&gt;
  
  
  Critical Assessment
&lt;/h3&gt;

&lt;p&gt;This appears to be a genuine capability leap, not pure marketing. The technical documentation demonstrates stepwise exploit development that goes well beyond what was previously possible with autonomous AI. The 4% to 85% increase in Firefox exploit success rate (per Anthropic's internal comparisons between Opus 4.6 and Mythos) is substantial.&lt;/p&gt;

&lt;p&gt;However, the &lt;em&gt;implications&lt;/em&gt; are where hype and reality diverge. The capability is real. Whether it necessitates the dramatic response Anthropic has mounted—and whether Anthropic is the appropriate custodian—is less clear.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Strategy: Controlled Release or Market Creation?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Anthropic's Stated Rationale
&lt;/h3&gt;

&lt;p&gt;Anthropic makes a straightforward argument: Frontier AI cybersecurity capabilities are approaching (or have reached) a level that could fundamentally alter the security landscape. By limiting Mythos Preview to vetted defensive partners, they give defenders time to harden systems before similar capabilities become broadly available to adversaries.&lt;/p&gt;

&lt;p&gt;This is framed as responsible AI governance—a model considered "too dangerous to release publicly" being deployed exclusively for defensive purposes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Business and Competitive Dimensions
&lt;/h3&gt;

&lt;p&gt;Forbes identifies five factors driving the invite-only rollout:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Real capability jump&lt;/strong&gt; (as discussed)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Responsible AI governance&lt;/strong&gt; positioning&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Strategic marketing through scarcity&lt;/strong&gt;—a narrative that generates enormous press&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capacity constraints&lt;/strong&gt;—Anthropic is throttling usage; the model is compute-intensive&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Premium pricing&lt;/strong&gt;—$25/$125 per million input/output tokens (versus $5/$25 for Opus), positioning Mythos as a luxury security product&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;VentureBeat adds crucial context: The same day Glasswing launched, Anthropic disclosed $30B in revenue and sealed the Google-Broadcom compute deal. The timing intersects with IPO speculation. A "high-profile, government-adjacent cybersecurity initiative with blue-chip partners is exactly the kind of program that burnishes an IPO narrative."&lt;/p&gt;

&lt;h3&gt;
  
  
  Who Actually Gains Access?
&lt;/h3&gt;

&lt;p&gt;The coalition structure creates an interesting dynamic. Tech competitors (Google vs. Microsoft) are both included. Smaller organizations and open-source maintainers are granted access via programs like "Claude for Open Source," with $4M in direct donations to open-source security organizations.&lt;/p&gt;

&lt;p&gt;But critics note this creates new forms of exclusion. As one Hacker News commenter observed: "The fact that you won't be able to produce secure software without access to one of these models. Good for them $."&lt;/p&gt;

&lt;p&gt;Whether the goal is truly defense for all, or defense for those who can afford/partner with Anthropic, is genuinely unclear.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Risks: Defense, Offense, and the Zero-Day Explosion
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The Core Paradox
&lt;/h3&gt;

&lt;p&gt;The fundamental challenge Mythos presents is that the same capabilities used by defenders to find and fix vulnerabilities can be used by attackers to find and exploit them. Anthropic acknowledges this explicitly but argues that "the advantage will belong to the side that can get the most out of these tools."&lt;/p&gt;

&lt;p&gt;In the short term, Anthropic warns, attackers who gain access to similar capabilities first could have a decisive advantage. In the long term, they expect defenders to prevail due to their ability to direct more resources and fix bugs before code ships.&lt;/p&gt;

&lt;p&gt;The "transitional period" could be tumultuous.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Happens When Adversaries Get Similar Models?
&lt;/h3&gt;

&lt;p&gt;Malware News reports serious concern within the intelligence community. Analysts are "casually chatting" about the Mythos release. Multiple officials note that U.S. agencies both defend networks and conduct offensive operations—and stockpile zero-days for future use.&lt;/p&gt;

&lt;p&gt;Hayden Smith of Hunted Labs calls the news "scary and ominous" because the offensive potential is unclear. "Even with deep vetting, the odds of Mythos flowing into the wrong hands is barely a hypothetical given the landscape of current attacks on the open source ecosystem."&lt;/p&gt;

&lt;p&gt;The concern isn't just state actors. As one executive at a cyber investment firm asked: "How is anyone supposed to defend against all of this at once?"&lt;/p&gt;

&lt;h3&gt;
  
  
  The Patching Problem
&lt;/h3&gt;

&lt;p&gt;Perhaps the most overlooked risk is the downstream impact of discovering thousands of vulnerabilities simultaneously. As Anthropic itself notes in its Red Team blog, "over 99% of the vulnerabilities we've found have not yet been patched."&lt;/p&gt;

&lt;p&gt;Flooding maintainers—many of whom are unpaid volunteers—with critical vulnerabilities at scale could overwhelm the very processes needed to fix them. Anthropic has built a triage pipeline to manually validate reports before submission, but bottlenecks seem inevitable.&lt;/p&gt;

&lt;p&gt;The 45-day coordinated disclosure window assumes maintainers can produce, test, and ship complex patches within that time—a presumption that may not hold for kernel-level vulnerabilities in critical systems.&lt;/p&gt;




&lt;h2&gt;
  
  
  Geopolitical Implications: AI as an Arms Race Component
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The U.S. Government Relationship
&lt;/h3&gt;

&lt;p&gt;Morgan Adamski, former executive director at U.S. Cyber Command, notes that "there's obviously a huge potential there from an adversarial perspective" for offensive use. She highlights an "equity conversation": if the U.S. exploits something in an adversarial network, it must also defend against that same vulnerability in its own infrastructure.&lt;/p&gt;

&lt;p&gt;Anthropic has briefed senior officials across the U.S. government on Mythos's capabilities, including both offensive and defensive applications. This comes after contentious disputes with the Pentagon over military uses of Claude, which saw Anthropic designated a "supply chain risk" before securing a preliminary injunction.&lt;/p&gt;

&lt;p&gt;Leah Siskind of the Foundation for Defense of Democracies argues: "The government 'needs to make amends with Anthropic and help them and Glasswing members maintain the American lead on AI by preventing Chinese model theft.'"&lt;/p&gt;

&lt;h3&gt;
  
  
  The International Dimension
&lt;/h3&gt;

&lt;p&gt;As Project Glasswing proceeds, other nations (particularly China, Russia, and U.S. adversaries) will almost certainly develop or acquire similar capabilities. Mythos-level models will eventually proliferate. The question isn't whether, but when—and whether the defensive advantages gained during the controlled rollout period will be durable.&lt;/p&gt;

&lt;p&gt;One concern: By making Mythos capabilities known while restricting access, Anthropic may have inadvertently created a roadmap for other AI labs to target. The technical specifications described in the system card provide a benchmark to aim for.&lt;/p&gt;




&lt;h2&gt;
  
  
  Trust and Irony: The Custodian Problem
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Anthropic's Security Track Record
&lt;/h3&gt;

&lt;p&gt;It is rich irony that Anthropic—asking governments and Fortune 500 companies to trust it with a model capable of autonomously exploiting Linux kernels—has suffered notable security lapses:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A draft Mythos blog post was left in an &lt;strong&gt;unsecured, publicly searchable data store&lt;/strong&gt; in March 2026, exposing roughly 3,000 internal assets&lt;/li&gt;
&lt;li&gt;For approximately three hours in March 2026, anyone running &lt;code&gt;npm install&lt;/code&gt; on Claude Code pulled down &lt;strong&gt;512,000 lines of Anthropic's source code&lt;/strong&gt; due to a packaging error&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Nicholas Carlini of Anthropic distinguishes these as "human errors in publishing tooling" rather than breaches of core security architecture—accurate as far as it goes, but a distinction that may not reassure stakeholders.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Boy Who Cried Wolf?
&lt;/h3&gt;

&lt;p&gt;There is legitimate concern about alarm fatigue. As Hacker News commenters note, every model is framed as revolutionizing everything, predicting doom if mishandled. When the next genuinely concerning capability arrives, will security practitioners—and the public—still be listening?&lt;/p&gt;

&lt;p&gt;Conversely, as others pointed out: "Tuning out completely because of the existence of false positives is not a good choice." The villagers may tire of the boy crying wolf, but wolves do eventually arrive.&lt;/p&gt;




&lt;h2&gt;
  
  
  Pros and Cons: A Critical Summary
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Pros
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Assessment&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Genuine capability improvement&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The demonstrated ability to autonomously find and chain vulnerabilities is a real step forward&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Proactive defense&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Finding bugs before adversaries do is fundamentally sound strategy&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Open-source support&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$4M in donations to OSS security addresses real asymmetries in resources&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Responsible disclosure pipeline&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Triage and human validation demonstrate awareness of maintenance bottlenecks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Transparency&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Detailed technical documentation with cryptographic commitments shows seriousness&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Coalition approach&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Bringing competitors together on security reduces fragmentation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Cons
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;Assessment&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Exclusionary access&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Creates dependency on Anthropic; smaller actors may be left behind&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;FOMO and coercion&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Organizations may join not out of belief but fear of seeming negligent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Overwhelmed maintainers&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Even with triage, the scale of findings risks swamping patching capacity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Verification limited&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Access restrictions make independent verification of claims difficult&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Business opportunism&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Timing with IPO and revenue milestones suggests mixed motives&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Geopolitical escalation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Demonstrating capabilities may accelerate adversarial AI development&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Trust issues&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Anthropic's security lapses undermine its credibility as gatekeeper&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Critical Opinions from Multiple Perspectives
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The Security Community
&lt;/h3&gt;

&lt;p&gt;On Hacker News, security professionals express a range of views:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Skeptical&lt;/strong&gt;: "This looks more like another lobby group...The 'urgency' is very likely mostly appreciated to drive policy."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Concerned&lt;/strong&gt;: "How is anyone supposed to defend against all of this at once?"&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measured&lt;/strong&gt;: "I side with you but on the other hand: this is how it works to get attention by those who aren't affiliated with computer science and AI."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Optimistic&lt;/strong&gt;: "At launch, a technology is considered dangerous for being too powerful. 3 months later, you are an absolute idiot to still be using that useless model."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Greg Kroah-Hartman's quote—about the "world switched" from AI slop to real reports—stands out as evidence from a respected figure in Linux development.&lt;/p&gt;

&lt;h3&gt;
  
  
  Industry Analysts
&lt;/h3&gt;

&lt;p&gt;Paulo Carvão at Forbes takes a nuanced view, noting both genuine capability and strategic positioning: "This announcement cannot be understood in isolation" from Anthropic's revenue growth and compute deals. The restricted rollout serves multiple purposes.&lt;/p&gt;

&lt;p&gt;Michael Nuñez at VentureBeat focuses on the fundamental wager: "Anthropic is, in essence, betting that transparency can outrun proliferation."&lt;/p&gt;

&lt;h3&gt;
  
  
  Intelligence and Government Concerns
&lt;/h3&gt;

&lt;p&gt;Morgan Adamski emphasizes the offense-defense equivalence: "If cyberintelligence analysts find a novel vulnerability in an enemy computer network, it's possible a U.S. system might have the same vulnerability, too."&lt;/p&gt;

&lt;p&gt;The intelligence community's "casual" discussions and serious concern about adversarial acquisition mirror the stakes: this isn't just a cybersecurity issue; it's a national security issue.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Open-Source Perspective
&lt;/h3&gt;

&lt;p&gt;Jim Zemlin, CEO of the Linux Foundation, provides perhaps the most compelling endorsement: "In the past, security expertise has been a luxury reserved for organizations with large security teams. Open-source maintainers—whose software underpins much of the world's critical infrastructure—have historically been left to figure out security on their own." Project Glasswing, he says, "offers a credible path to changing that equation."&lt;/p&gt;

&lt;p&gt;This gets at a real problem: the asymmetry between well-resourced corporations and the volunteer-maintained projects that form software's foundation.&lt;/p&gt;




&lt;h2&gt;
  
  
  Conclusion: A Necessary Step, But A Flawed One?
&lt;/h2&gt;

&lt;p&gt;Project Glasswing represents a genuinely significant moment in AI development. The technical capabilities of Claude Mythos Preview appear real enough that Anthropic—not a company known for understatement—is willing to frame them as too dangerous for public release. The decision to limit access to defensive partners and invest in open-source security is, in principle, defensible.&lt;/p&gt;

&lt;p&gt;But the initiative is also deeply problematic:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;It concentrates power&lt;/strong&gt; in Anthropic's hands during a transition period that will be contested globally&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It markets through scarcity&lt;/strong&gt;, creating artificial urgency that serves business interests&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It may overwhelm&lt;/strong&gt; the very maintenance processes needed to address discovered vulnerabilities&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It invites escalation&lt;/strong&gt;, as other labs rush to match or exceed demonstrated capabilities&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It suffers from trust deficits&lt;/strong&gt;, given Anthropic's own security history and the incentives of a company on an IPO trajectory&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The core question—whether Project Glasswing genuinely makes the world more secure, or merely reshapes advantage within existing power structures—has no clear answer yet. The only certainty is that the age of AI-augmented cyberconflict has begun in earnest. The glasswing's transparent wings hide vulnerabilities well. But in seeking to reveal those vulnerabilities to defenders first, Anthropic may have revealed something else: just how quickly the ground beneath cybersecurity's feet is shifting.&lt;/p&gt;

&lt;p&gt;In the coming months—before the next frontier lab announces its own game-changing model, before adversarial access reaches Mythos-equivalent levels, before the inevitable disclosure of vulnerabilities that even Anthropic cannot contain—we will learn whether controlled releases like Project Glasswing can genuinely preserve a defensive advantage, or whether the fundamental symmetries of offense and defense make this a game of diminishing returns.&lt;/p&gt;

&lt;p&gt;The wolf may or may not have arrived. But when it does, the villages that invested in defenses during the calm will have a better chance. Whether Anthropic should be the one selling those defenses is the question that remains.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>cybersecurity</category>
      <category>news</category>
    </item>
    <item>
      <title>Beyond OpenClaw: The Rise of the Lightweight AI Agent Ecosystem in 2026</title>
      <dc:creator>TheRabbitHole</dc:creator>
      <pubDate>Fri, 06 Mar 2026 10:43:54 +0000</pubDate>
      <link>https://dev.to/therabbithole/beyond-openclaw-the-rise-of-the-lightweight-ai-agent-ecosystem-in-2026-j91</link>
      <guid>https://dev.to/therabbithole/beyond-openclaw-the-rise-of-the-lightweight-ai-agent-ecosystem-in-2026-j91</guid>
      <description>&lt;p&gt;OpenClaw (originally Clawdbot) has long been the dominant force in autonomous AI agents, boasting over 267,000 GitHub stars. But as its codebase has ballooned to over 430,000 lines, developers have begun to voice concerns over its massive resource footprint and security vulnerabilities.&lt;/p&gt;

&lt;p&gt;In response, a "small-is-beautiful" revolution has taken over GitHub. Developers are flocking to lightweight, transparent alternatives that prioritize security, auditability, and efficiency.&lt;/p&gt;

&lt;p&gt;If you are looking for projects similar to &lt;strong&gt;NanoClaw&lt;/strong&gt;, here is your comprehensive guide to the ecosystem of lightweight alternatives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Top Open-Source Lightweight Alternatives
&lt;/h2&gt;

&lt;p&gt;These projects share a common philosophy: a smaller codebase means better auditability and lower resource usage.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. NanoClaw
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Language:&lt;/strong&gt; TypeScript (Node.js)&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;GitHub Stars:&lt;/strong&gt; ~19,500&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Focus:&lt;/strong&gt; Security-First &amp;amp; Container Isolation&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;The Pitch:&lt;/strong&gt; NanoClaw is the go-to choice for security-conscious developers. Unlike the original OpenClaw, which often runs in a single process with shared memory, NanoClaw forces &lt;strong&gt;OS-level container isolation&lt;/strong&gt; (e.g., Apple Containers on macOS). This ensures that even if an agent goes rogue, it cannot access your host machine's filesystem or sensitive &lt;code&gt;.env&lt;/code&gt; credentials. It integrates seamlessly with the Claude Code ecosystem.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Nanobot (University of Hong Kong)
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Language:&lt;/strong&gt; Python&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;GitHub Stars:&lt;/strong&gt; ~29,400&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Focus:&lt;/strong&gt; Extreme Transparency &amp;amp; Simplicity&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;The Pitch:&lt;/strong&gt; If your goal is to learn or customize, Nanobot is unmatched. It is roughly &lt;strong&gt;4,000 lines of Python&lt;/strong&gt;—about 99% smaller than OpenClaw. Despite its tiny size, it packs in persistent memory, web search, and integrations for Telegram and WhatsApp.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. ZeroClaw
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Language:&lt;/strong&gt; Rust&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;GitHub Stars:&lt;/strong&gt; ~23,700&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Focus:&lt;/strong&gt; High Performance &amp;amp; Safety&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;The Pitch:&lt;/strong&gt; For the production environment, ZeroClaw offers "Claw done right." It compiles down to a &lt;strong&gt;3.4 MB binary&lt;/strong&gt; and uses less than 5 MB of RAM at runtime. Its standout feature is being "secure-by-default" with strict workspace scoping for filesystems.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4. NullClaw
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Language:&lt;/strong&gt; Zig&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;GitHub Stars:&lt;/strong&gt; ~5,480&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Focus:&lt;/strong&gt; Ultra-Minimalist Runtime&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;The Pitch:&lt;/strong&gt; NullClaw is extreme minimalism incarnate. It produces a static binary of only ~678 KB that boots in milliseconds. It is the ideal candidate for edge devices and IoT scenarios where every byte counts.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5. PicoClaw
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Language:&lt;/strong&gt; Go&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;GitHub Stars:&lt;/strong&gt; ~12,000+&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Focus:&lt;/strong&gt; Embedded Hardware &amp;amp; IoT&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;The Pitch:&lt;/strong&gt; PicoClaw is designed to run on cheap hardware. It can operate on &lt;strong&gt;$10 RISC-V boards&lt;/strong&gt; with less than 10 MB of RAM. It also includes free voice transcription via Groq Whisper, making it a powerhouse for hobbyists working on embedded projects.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Summary Comparison
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Language&lt;/th&gt;
&lt;th&gt;Footprint&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;NanoClaw&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Node.js&lt;/td&gt;
&lt;td&gt;Small&lt;/td&gt;
&lt;td&gt;Security-first / Container isolation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Nanobot&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Python&lt;/td&gt;
&lt;td&gt;~4K lines&lt;/td&gt;
&lt;td&gt;Learning / Simple customization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ZeroClaw&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Rust&lt;/td&gt;
&lt;td&gt;&amp;lt;5 MB RAM&lt;/td&gt;
&lt;td&gt;High performance / Safety&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;NullClaw&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Zig&lt;/td&gt;
&lt;td&gt;678 KB&lt;/td&gt;
&lt;td&gt;Extreme edge/IoT minimalism&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;PicoClaw&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Go&lt;/td&gt;
&lt;td&gt;&amp;lt;10 MB RAM&lt;/td&gt;
&lt;td&gt;Cheap embedded hardware&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Specialized &amp;amp; Enterprise Alternatives
&lt;/h2&gt;

&lt;p&gt;While the projects above focus on being lightweight, other alternatives are targeting specific enterprise niches:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;memU:&lt;/strong&gt; Focuses on "proactive" assistance using a Hierarchical Knowledge Graph for superior long-term memory.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Moltworker:&lt;/strong&gt; A serverless version of OpenClaw hosted on Cloudflare Workers, offering sandboxed execution without local machine access.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Adopt AI:&lt;/strong&gt; An enterprise-grade platform that automates API discovery and action generation for complex corporate workflows.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;TinyClaw:&lt;/strong&gt; A multi-agent system that coordinates specialized agents (coder, researcher, etc.) in parallel via a live terminal dashboard.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Which Projects Are Rising the Fastest?
&lt;/h2&gt;

&lt;p&gt;As of March 2026, the growth charts show a clear divide between the established educational tools and the new production-ready contenders.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;🚀 PicoClaw (The Viral Leader):&lt;/strong&gt; Gained over 12,000 stars in its first week. Its ability to run on $10 hardware has captivated the maker community.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;📈 ZeroClaw (The Pro Choice):&lt;/strong&gt; Seeing a surge in professional adoption. It is currently the preferred choice for developers wanting a robust, "agentic OS" workflow.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;🛡️ NanoClaw (The Security Pick):&lt;/strong&gt; Growing rapidly among security circles, particularly due to its recent "Agent Swarms" update and compatibility with Claude Code.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to Choose?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Choose Nanobot&lt;/strong&gt; if you want to read the code and understand how agents work.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Choose ZeroClaw&lt;/strong&gt; if you need speed and memory safety for a production app.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Choose NanoClaw&lt;/strong&gt; if you are handling sensitive data and need strict container isolation.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Choose PicoClaw&lt;/strong&gt; if you want to build AI into physical devices on a budget.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The era of the "bloated agent" is ending. With tools like these, the future of autonomous AI is fast, secure, and accessible.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Agents Can Now Clone Themselves and Do Crazy Things (Part I: Deep Stock Analysis)</title>
      <dc:creator>TheRabbitHole</dc:creator>
      <pubDate>Mon, 26 Jan 2026 10:17:48 +0000</pubDate>
      <link>https://dev.to/therabbithole/agents-can-now-clone-themselves-and-do-crazy-things-part-i-deep-stock-analysis-71l</link>
      <guid>https://dev.to/therabbithole/agents-can-now-clone-themselves-and-do-crazy-things-part-i-deep-stock-analysis-71l</guid>
      <description>&lt;p&gt;Most chatbots, such as ChatGPT and Claude, are becoming more powerful every day. They are incorporating more tools, characters and features, such as Canvas or Artifacts, to improve usability. However, especially if you are a heavy user of AI (especially as a non-coder), the limitations are the same: the more data and the more complex the tasks, the less AI becomes usable.&lt;/p&gt;

&lt;p&gt;It becomes lazy and takes shortcuts.&lt;/p&gt;

&lt;p&gt;It hallucinates. It forgets things. The quality degrades massively, and worst of all, you still pay for it.&lt;/p&gt;

&lt;p&gt;Most of these issues are known limitations that happen because of one of the most limiting factors of AI: the context window. Think of it as the AI's limited working memory: the more data it contains, the more overwhelmed the AI becomes while still trying to please you. The result is a pure waste of time and money.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Solution That Changes Everything
&lt;/h2&gt;

&lt;p&gt;There have been a lot of advancements in this area trying to overcome these technical limitations, such as plugging in memories, but one incredibly powerful solution is multi-agency.&lt;/p&gt;

&lt;p&gt;The AI breaks down tasks it has never seen before using its reasoning capabilities and sends them to other AIs (so-called subagents) to complete. Then it aggregates the results and answers the user's request.&lt;/p&gt;

&lt;p&gt;In this approach, the so-called sub-agent starts with a fresh memory. It doesn't need to know the entire context; it just needs to know the subtask at hand. It executes the task, delivers the results and disappears. Any further subtasks start with a new LLM. This core difference to having one large LLM trying to do everything by itself changes the entire game.&lt;/p&gt;

&lt;p&gt;Handling much more complex tasks becomes possible. You get much less hallucination and much higher quality. Think of those subagents focusing on one smaller task; they can perform much better than trying to handle a huge task all at once. And if you have parallelisation, the end-to-end experience can be much faster than single processing, though this also depends on the tooling of the multi-agent solution.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Tools You Can Use Right Now
&lt;/h2&gt;

&lt;p&gt;If you follow the news, you might have heard about Claude Cowork. Built on top of a framework developed by Anthropic a few months ago, called Agent SDK, Claude Cowork can process highly complex tasks end-to-end using a high-reasoning, multi-agent approach.&lt;/p&gt;

&lt;p&gt;It develops a well-thought-out plan for accomplishing a given complex task from start to finish. It spawns multiple agents ad hoc (think of it as a scalable team on demand). It extends code in a sandbox environment, giving users the full power of coding without requiring any prior knowledge (e.g. reading and editing files, calling APIs, and much more).&lt;/p&gt;

&lt;p&gt;This tool is incredibly powerful, but expensive, though worth the investment if you consider the ROI.&lt;/p&gt;

&lt;p&gt;If you are reluctant to pay a monthly subscription fee of $100 to $200, you can also use the framework with code, or you can use Cherry Studio, an open-source chatbot that integrates this framework.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F790zbu5iov8yyw4t6oeg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F790zbu5iov8yyw4t6oeg.png" alt=" " width="800" height="496"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A Real World Example: Deep Analysis of Microsoft's 2025 Annual Report
&lt;/h2&gt;

&lt;p&gt;This technology can be used to solve a variety of complex tasks, including those that require the use of tools. Imagine presenting a dense financial report to different experts (financial gurus, strategists, etc.) to obtain a comprehensive view of the results.&lt;/p&gt;

&lt;p&gt;The coordinating AI (the one you are talking to in the chat) decides ad hoc how many agents to use, how to prompt them, and so on. You don't need any prior configuration. That's the real beauty of this amazing technology.&lt;/p&gt;

&lt;p&gt;The process works like this: First, the system reads the contents of the report, then sends subtasks to multiple expert subagents. Each of these subtasks is a subagent with its own memory and tools. After a minute or so, you have a detailed analysis of the final report compiled from five different angles.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost Considerations
&lt;/h2&gt;

&lt;p&gt;You might be wondering how much this will cost you. For a dense report with millions of tokens processed, you're looking at roughly $2.50 to $3.00 USD using Haiku 4.5, especially when cached tokens reduce the total cost significantly. If there's a lot at stake for you, it's more than worth every penny.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting Started in Three Steps
&lt;/h2&gt;

&lt;p&gt;Try it yourself with Cherry Studio. Install Cherry Studio from the official repository, add the API key for Anthropic, and click 'Add Agent' on the right. Then select the model and create a scratch area. That's it.&lt;/p&gt;

&lt;p&gt;Now you can start chatting with the agent and let it free you from those painful, boring tasks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Read the full deep dive on airabbit.blog:&lt;/strong&gt; &lt;a href="https://airabbit.blog/agents-can-now-clone-themselves-and-do-crazy-things-part-i-deep-stock-analysis/" rel="noopener noreferrer"&gt;https://airabbit.blog/agents-can-now-clone-themselves-and-do-crazy-things-part-i-deep-stock-analysis/&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Is The Future of AI is On-Demand?</title>
      <dc:creator>TheRabbitHole</dc:creator>
      <pubDate>Sat, 24 Jan 2026 19:11:21 +0000</pubDate>
      <link>https://dev.to/therabbithole/the-future-of-ai-is-on-demand-4a73</link>
      <guid>https://dev.to/therabbithole/the-future-of-ai-is-on-demand-4a73</guid>
      <description>&lt;p&gt;Recently, a friend of mine who has no affiliation with IT whatsoever approached me with great excitement about an app he had developed overnight. He built the whole thing on his phone. I was baffled, though not surprised. These days, almost anything is possible — or at least, we like to think so.&lt;/p&gt;

&lt;p&gt;This new reality makes technology accessible to almost everyone. All you need is an idea, a phone and a subscription for a month or so, and you're good to go, right?&lt;/p&gt;

&lt;p&gt;Forgetting for a moment the 'crimes' that laypeople are committing regarding day-two operations (patching, security, etc.), the world is already flooded with apps. Everyone has their own business model, subscription process and requirements for signing up.&lt;/p&gt;

&lt;p&gt;For consumers, this is becoming a nightmare.&lt;/p&gt;

&lt;p&gt;Sharing your personal data with each and every one of them.&lt;/p&gt;

&lt;p&gt;Paying everyone a subscription.&lt;/p&gt;

&lt;p&gt;And so on.&lt;/p&gt;

&lt;p&gt;I used to have lots of these apps and subscriptions one or two years ago.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Presentation AI&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;AI chatbots (Claude, ChatGPT, etc.).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Canva&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;AI video and image generators (Runway, etc.).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Freepik&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And many more.&lt;/p&gt;

&lt;p&gt;And that’s just for AI!&lt;/p&gt;

&lt;p&gt;I have started to cancel a lot of subscriptions, including ChatGPT and Claude. I have started switching to platforms that aggregate all of these solutions in one place, with one account and one subscription — and that’s it! &lt;/p&gt;

&lt;p&gt;This has shown me that I don't actually need to pay for a monthly or yearly subscription just to generate ads (like AdCreative) or flyers (like Canva). I do a lot, but I don't need a permanent subscription for that.&lt;/p&gt;

&lt;h3&gt;
  
  
  How Aggregation Platforms Work
&lt;/h3&gt;

&lt;p&gt;Aggregation platforms such as Poe and Apify — and, I believe, ChatGPT in the future — bring together all the services and apps available. Think of it as a 'pay once, use all' model, with the amount depending on the subscription plan. &lt;/p&gt;

&lt;p&gt;This is different from Amazon, where you just have a directory and pay each one individually (this is what we have now).&lt;/p&gt;

&lt;p&gt;Apify is one amazing platform that has proven how powerful this business model is.&lt;/p&gt;

&lt;p&gt;When you subscribe to Apify, you get access to around 5,000 "actors", most of which have flexible pricing options, such as paying per output result or even per call.&lt;/p&gt;

&lt;p&gt;For example, I pay $50 per month and can use:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;LinkedIn actors to scrape LinkedIn;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Reddit actors to scrape Reddit.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;data analytics actors, such as Semrush, for in-depth analysis.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;and many more&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With pay-per-use, I don't have to pay for the Reddit API or a Semrush subscription. You get my point.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Future of Aggregation
&lt;/h3&gt;

&lt;p&gt;Now, think of this same concept with ChatGPT Store. We could have these giant platforms hosting thousands of AI services for everything:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Creative writing&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Generating presentations&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Generating images&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Or even entire videos or books.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And all on a pay-per-use basis. This is technically already possible but still at a very early stage.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Caveat
&lt;/h3&gt;

&lt;p&gt;One could think of monopoly platforms such as Amazon, and of course, serious concerns arise with regard to control, power and security. However, we must also consider:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;How much power do they exert?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;How do they monetise developers?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The policy: what does and doesn't match their strategy.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;A single point of failure.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In an ideal world, there would be multiple platforms that aggregate services, eliminating the need for multiple registrations and payments, and saving time and money on testing things that we rarely use — and even worse, things that don’t fulfil their promises, which we often only realise after paying a hefty subscription.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Stop Trying to Pick the 'Best' LLM. Let Them Answer Together (For Under a Dime)</title>
      <dc:creator>TheRabbitHole</dc:creator>
      <pubDate>Sat, 24 Jan 2026 16:48:22 +0000</pubDate>
      <link>https://dev.to/therabbithole/stop-trying-to-pick-the-best-llm-let-them-answer-together-for-under-a-dime-12pl</link>
      <guid>https://dev.to/therabbithole/stop-trying-to-pick-the-best-llm-let-them-answer-together-for-under-a-dime-12pl</guid>
      <description>&lt;p&gt;We've all been there. You ask ChatGPT for architectural advice, and it gives you a confident answer. But something nags at you — is this actually the best approach, or just the first one the model latched onto?&lt;/p&gt;

&lt;p&gt;Single models have blind spots. They're trained on specific datasets, optimized for certain response patterns, and prone to confident-but-wrong answers. Getting a second opinion from a different model helps, but manually copying prompts between interfaces is tedious.&lt;/p&gt;

&lt;p&gt;What if you could query multiple top-tier models simultaneously and see where they agree, disagree, or bring up angles you hadn't considered?&lt;/p&gt;

&lt;p&gt;That's exactly what &lt;strong&gt;Super AI Bench&lt;/strong&gt; does. It's an MCP (Model Context Protocol) server that acts as your AI consensus engine, automatically querying the smartest available models and synthesizing their responses.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Simple Idea: AI as a Panel, Not an Oracle
&lt;/h2&gt;

&lt;p&gt;Instead of treating AI as a single expert, think of it as a panel of specialists. Each model has different training data, architecture, and "experience":&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude&lt;/strong&gt; tends toward careful, nuanced analysis with strong ethical considerations&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPT-4&lt;/strong&gt; excels at structured reasoning and technical implementation details
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini&lt;/strong&gt; often brings in creative angles and cross-domain connections&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mistral&lt;/strong&gt; might prioritize efficiency and practical constraints&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When they converge on an answer, you can be more confident. When they diverge, you see the complexity instead of getting a false sense of certainty.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real Example: Debugging a Production Issue
&lt;/h2&gt;

&lt;p&gt;Let's say you're troubleshooting a memory leak. Here's what a multi-model consensus looks like in practice:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your prompt:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"Node.js app memory grows 2% hourly. Heap dumps show string accumulation. 
Using Express, Redis, and Winston. Where should I look?"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Consensus results:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"models_queried"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"response_time"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"8.3s"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"consensus"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"high_confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Check Winston transport configuration"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Review Redis connection string handling"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Look for unclosed response streams"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"divergent_opinions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"claude_3.5"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Mentioned event listener leaks in error handlers specifically"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"gpt_4"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Suggested checking for large request/response logging"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"gemini_1.5"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Flagged potential issues with custom formatters retaining references"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"unique_insights"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"One model spotted that your Redis retry strategy might be buffering commands"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Another noted that Winston's FileTransport with high logging levels can accumulate"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fnc34lggo4no42gsy5n4w.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fnc34lggo4no42gsy5n4w.png" alt=" " width="800" height="190"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Instead of one model's best guess, you get a prioritized checklist and discover edge cases you might have missed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Strategic Example: Business Decision Making
&lt;/h2&gt;

&lt;p&gt;Imagine you're a product manager deciding whether to pivot your SaaS platform toward AI features or double down on core functionality.&lt;/p&gt;

&lt;p&gt;This isn't a technical question. It's strategic, involves market assumptions, financial risk, and competitive positioning. A single AI model will give you &lt;em&gt;one perspective&lt;/em&gt; with high confidence. But what are you missing?&lt;/p&gt;

&lt;p&gt;With Super AI Bench, you send one prompt: &lt;em&gt;"Our SaaS has 5K users, strong retention, but slower feature velocity than competitors. Should we pivot to add AI features or strengthen core product? Consider: market timing, engineering cost, user retention risk, competitive threat."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What you get back:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude&lt;/strong&gt; focuses on user risk and thoughtful long-term strategy ("Don't chase trends; validate demand first")&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPT-4&lt;/strong&gt; brings structured business analysis ("Calculate CAC impact on both paths; model the revenue upside")&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini&lt;/strong&gt; surfaces market dynamics you hadn't considered ("AI features become table stakes in 12 months for your category")&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mistral&lt;/strong&gt; emphasizes resource constraints ("You don't have the engineering bandwidth for both")&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Instead of one confident answer, you see the trade-offs clearly. You discover that the real decision isn't "pivot or not" — it's "whether you have the team capacity to do it well." That insight alone might save you six months of wasted effort.&lt;/p&gt;

&lt;p&gt;This is where consensus becomes valuable: not because the models are always right, but because you see the problem from multiple angles instead of getting a false sense of certainty from a single perspective.&lt;/p&gt;

&lt;h2&gt;
  
  
  More Affordable Than You Might Think
&lt;/h2&gt;

&lt;p&gt;Running multiple models sounds expensive, but for many use cases, the cost is surprisingly low. Most queries cost less than a penny, and even complex analyses rarely exceed a few cents.&lt;/p&gt;

&lt;p&gt;Here are a few real examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Quick technical question&lt;/strong&gt;: 3 models responded in under 1 second total, cost was less than $0.01&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Detailed code review&lt;/strong&gt;: 3 models took 7-34 seconds, cost was $0.01-$0.02&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Complex architecture discussion&lt;/strong&gt;: Multiple models provided detailed responses for less than $0.02 total&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When you consider the cost of a wrong decision or missed bug, spending a few cents to get multiple perspectives is a pragmatic investment.&lt;/p&gt;

&lt;h2&gt;
  
  
  When This Actually Helps
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;✅ Good use cases:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;High-stakes decisions&lt;/strong&gt; where blind spots are costly (architecture, security)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Creative blocks&lt;/strong&gt; when you need fresh perspectives (marketing campaigns, product features)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Risk assessment&lt;/strong&gt; to surface concerns you hadn't considered&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Learning complex topics&lt;/strong&gt; by seeing different explanation styles&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fact-checking&lt;/strong&gt; controversial claims by checking for consensus&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;❌ Don't bother when:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You need a quick, simple answer ("What's the Python string length function?")&lt;/li&gt;
&lt;li&gt;The task is deterministic (math calculations, code syntax)&lt;/li&gt;
&lt;li&gt;You're on a tight budget (5 models = 5x the API costs)&lt;/li&gt;
&lt;li&gt;You already have deep expertise in the domain&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Honest Limitations
&lt;/h2&gt;

&lt;p&gt;This isn't magic. It's pattern matching at scale.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cost&lt;/strong&gt;: Running 5 top-tier models isn't cheap. Use it for important questions, not every query.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Speed&lt;/strong&gt;: You'll wait 5-10 seconds for all responses. It's not for real-time applications.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agreement ≠ Truth&lt;/strong&gt;: Models can all be wrong in the same direction. They share some training data and architectural biases.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Divergence ≠ Uselessness&lt;/strong&gt;: Sometimes the outlier model catches something critical. The "consensus" is just a starting point for your own judgment.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Not Just Another Multi-Model Chatbot
&lt;/h2&gt;

&lt;p&gt;You might be thinking: "Can't I just use one of those open-source chatbots that let me select multiple models and send them the same prompt?"&lt;/p&gt;

&lt;p&gt;This is fundamentally different.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Open-source multi-model chatbots are static&lt;/strong&gt; - You have to manually choose which models to query, copy your prompt to each one, and then manually compare the responses yourself. It's a tedious, repetitive process that doesn't scale.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Super AI Bench is dynamic and AI-driven&lt;/strong&gt; - The AI assistant frames your question, automatically determines which models are most suitable based on live benchmarks, sends the prompt to them in parallel, and aggregates the results into a coherent summary. All without any interaction from you after the initial prompt.&lt;/p&gt;

&lt;p&gt;The difference is night and day:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Before&lt;/strong&gt;: "Let me check 3 different models manually..." &lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;After&lt;/strong&gt;: "Hey AI, what's the best approach here?" (30 seconds, fully automated)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This isn't just about querying multiple models - it's about intelligent orchestration that removes the friction entirely.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup in 30 Seconds
&lt;/h2&gt;

&lt;p&gt;Getting started is simpler than you might think. You only need two accounts:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Apify account&lt;/strong&gt; - Free tier available, and login uses OAuth (no password needed)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Replicate account&lt;/strong&gt; - For accessing the AI models, just grab your API key&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That's it. No complex configuration, no infrastructure to manage.&lt;/p&gt;

&lt;p&gt;Add this to your MCP settings:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"super-ai-bench"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"npx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"-y"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"mcp-remote"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"https://flamboyant-leaf--super-ai-bench-mcp.apify.actor/mcp?replicateApiKey=&amp;lt;REPLICATE_API_KEY&amp;gt;"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Just replace &lt;code&gt;&amp;lt;REPLICATE_API_KEY&amp;gt;&lt;/code&gt; with your actual key. Apify handles authentication automatically through OAuth when you first use the actor.&lt;/p&gt;

&lt;p&gt;From that point forward, simply select the "Super AI Bench" MCP in your AI assistant, frame your question, and let it query multiple models and summarize the responses for you. The actor manages all the parallel calls, error handling, and response formatting.&lt;/p&gt;

&lt;p&gt;See the README for more configuration options and advanced usage patterns.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Testing MCP Servers like a Pro using MCPJam Inspector</title>
      <dc:creator>TheRabbitHole</dc:creator>
      <pubDate>Wed, 21 Jan 2026 11:37:07 +0000</pubDate>
      <link>https://dev.to/therabbithole/testing-mcp-servers-like-a-pro-using-mcpjam-inspector-22ka</link>
      <guid>https://dev.to/therabbithole/testing-mcp-servers-like-a-pro-using-mcpjam-inspector-22ka</guid>
      <description>&lt;p&gt;Building and testing MCP (Model Context Protocol) servers is frustrating without the right tools. Most developers waste hours switching between different environments—writing code, then switching to clients like Cursor or Claude Desktop just to test a simple function call, then back to the IDE to debug issues. You're constantly guessing what's wrong when tools fail: Is it the MCP protocol implementation? The connection parameters? The tool definition? MCPJam Inspector solves this by giving you a dedicated, visual workspace for testing, debugging, and validating MCP servers without ever leaving your development flow. It's the difference between fumbling in the dark and having X-ray vision into your MCP implementation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is MCPJam Inspector?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;MCPJam Inspector&lt;/strong&gt; is a local-first developer tool for testing, debugging, and inspecting Model Context Protocol (MCP) servers and ChatGPT/OpenAI apps. Think of it as "Postman for MCP"—a visual interface that lets you explore, test, and debug MCP servers without needing to deploy them or connect through production clients.&lt;/p&gt;

&lt;h3&gt;
  
  
  Key Features
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Visual Server Management&lt;/strong&gt;: Connect to MCP servers via STDIO, HTTP, or SSE protocols&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool Testing&lt;/strong&gt;: Manually invoke and test MCP tools with custom parameters&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Resource Inspection&lt;/strong&gt;: Browse and fetch resources exposed by MCP servers&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompt Templates&lt;/strong&gt;: Test and use prompt templates with slash commands&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LLM Playground&lt;/strong&gt;: Simulate how your MCP server performs with various LLMs (OpenAI, Claude, Ollama, etc.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real-time Logging&lt;/strong&gt;: View all JSON-RPC messages, requests, responses, and errors&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OAuth Debugging&lt;/strong&gt;: Test and debug OAuth flows for authenticated servers&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chat Interface&lt;/strong&gt;: Interact with your MCP server conversationally&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  What Can You Use It For?
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Development&lt;/strong&gt;: Build and test MCP-based tools locally without switching to clients like Cursor or Claude Desktop&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;QA &amp;amp; Debugging&lt;/strong&gt;: Validate tool definitions, prompt templates, and resource endpoints against the MCP specification&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Experimentation&lt;/strong&gt;: Test your MCP server with different LLM models to see how it behaves in various contexts&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Learning&lt;/strong&gt;: Understand how MCP servers work by inspecting the protocol messages in real-time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Integration Testing&lt;/strong&gt;: Verify that your MCP server works correctly before deploying to production&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  About This Tutorial
&lt;/h2&gt;

&lt;p&gt;This tutorial demonstrates how to use &lt;strong&gt;MCPJam Inspector&lt;/strong&gt; to add and test MCP servers. We use the Tavily MCP server as an example, but &lt;strong&gt;the same process works for any MCP server&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Custom MCP servers you've built&lt;/li&gt;
&lt;li&gt;Third-party MCP servers (GitHub, Slack, Notion, etc.)&lt;/li&gt;
&lt;li&gt;Local MCP servers running on your machine&lt;/li&gt;
&lt;li&gt;Remote MCP servers via HTTP/SSE&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The steps are identical—just replace the server URL and configuration with your own MCP server details.&lt;/p&gt;

&lt;h2&gt;
  
  
  Overview
&lt;/h2&gt;

&lt;p&gt;This tutorial walks you through using MCPJam Inspector to:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Add an MCP server (using Tavily as an example)&lt;/li&gt;
&lt;li&gt;Connect via HTTP/SSE&lt;/li&gt;
&lt;li&gt;View available tools from the server&lt;/li&gt;
&lt;li&gt;Test the tools with custom parameters&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;MCPJam Inspector running at &lt;code&gt;http://127.0.0.1:6274&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;An MCP server to connect to (we'll use Tavily as an example - get an API key from &lt;a href="https://tavily.com" rel="noopener noreferrer"&gt;Tavily's website&lt;/a&gt; if following along)&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Step 1: Open MCPJam Inspector
&lt;/h2&gt;

&lt;p&gt;Navigate to &lt;code&gt;http://127.0.0.1:6274&lt;/code&gt; in your browser. You'll see the main dashboard with no servers connected.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4z79gh2fog80tyorjm91.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F4z79gh2fog80tyorjm91.png" alt="Initial State" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Click "Add Server"
&lt;/h2&gt;

&lt;p&gt;Click the &lt;strong&gt;"Add Server"&lt;/strong&gt; button in the top right corner of the MCP Servers section.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9nri75jrfvutbspjztsb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F9nri75jrfvutbspjztsb.png" alt="Add Server Dialog" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: Select HTTP Connection Type
&lt;/h2&gt;

&lt;p&gt;The dialog opens with STDIO selected by default. Click the &lt;strong&gt;Connection Type&lt;/strong&gt; dropdown and select &lt;strong&gt;"HTTP"&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F28kh1zelo41o588gb1dp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F28kh1zelo41o588gb1dp.png" alt="Connection Type Dropdown" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After selecting HTTP, the form changes to show HTTP-specific fields:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Server Name&lt;/strong&gt;: Enter a name for your server&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;URL&lt;/strong&gt;: Enter the Tavily MCP server URL&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Authentication&lt;/strong&gt;: Configure if needed&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Custom Headers&lt;/strong&gt;: Add any custom headers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fywjhfu6bq53fn0jgfbcg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fywjhfu6bq53fn0jgfbcg.png" alt="HTTP Connection Form" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Server Name&lt;/strong&gt;: Enter a name for your server (we used &lt;code&gt;tavily&lt;/code&gt; as an example)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;URL&lt;/strong&gt;: Enter your MCP server URL. For the Tavily example:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;   https://mcp.tavily.com/mcp/?tavilyApiKey=YOUR_API_KEY
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Note&lt;/strong&gt;: Replace &lt;code&gt;YOUR_API_KEY&lt;/code&gt; with your actual API key. For other MCP servers, use their respective connection URLs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fgo0rl0okp6byzmnckdhr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fgo0rl0okp6byzmnckdhr.png" alt="Server Name Filled" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fnneg63xy0ujmsy31ib9n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fnneg63xy0ujmsy31ib9n.png" alt="URL Filled" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Click the &lt;strong&gt;"Add Server"&lt;/strong&gt; button at the bottom of the dialog. The server will connect automatically.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbcda53hbokpckg6bqxda.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbcda53hbokpckg6bqxda.png" alt="Server Connected" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You can see:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Server name: &lt;strong&gt;tavily&lt;/strong&gt; (or whatever you named it)&lt;/li&gt;
&lt;li&gt;Connection type: &lt;strong&gt;HTTP/SSE&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Status: &lt;strong&gt;Connected&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Server version: &lt;strong&gt;v2.14.2&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Click on &lt;strong&gt;"Tools"&lt;/strong&gt; in the left sidebar to see all available tools from your MCP server.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F1rua2w520z4v8o64sbun.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F1rua2w520z4v8o64sbun.png" alt="Tools List" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In our example with Tavily, we see 4 tools:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;tavily_search&lt;/strong&gt; - Search the web for real-time information&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;tavily_extract&lt;/strong&gt; - Extract content from specific web pages&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;tavily_crawl&lt;/strong&gt; - Crawl multiple pages from a website&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;tavily_map&lt;/strong&gt; - Map and discover website structure&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Different MCP servers will expose different tools based on their functionality.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 8: Test a Tool
&lt;/h2&gt;

&lt;p&gt;Click on any tool from your MCP server to open its configuration form. In our example, we'll test &lt;strong&gt;"tavily_search"&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdxbp2fmpmbveaszeztsb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdxbp2fmpmbveaszeztsb.png" alt="Tool Form" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The form shows all available parameters for the selected tool. Each MCP server's tools will have different parameters based on their functionality.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 9: Enter Parameters
&lt;/h2&gt;

&lt;p&gt;Fill in the required parameters. For the tavily_search example, enter a test query like: &lt;code&gt;MCP protocol tutorial&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2em42xw6g0d5cy73kscq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2em42xw6g0d5cy73kscq.png" alt="Query Filled" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 10: Execute the Tool
&lt;/h2&gt;

&lt;p&gt;Click the &lt;strong&gt;"Execute"&lt;/strong&gt; button to run the tool. The button will show "Running" while processing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fe6wt7rgk2a6ijycka87a.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fe6wt7rgk2a6ijycka87a.png" alt="Search Results" width="800" height="474"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The results will appear in the &lt;strong&gt;Response&lt;/strong&gt; section below, showing the tool's output in a structured format. The exact format depends on what the tool returns.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next Steps
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Explore More Tools
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Try other tools from your MCP server&lt;/li&gt;
&lt;li&gt;Test different parameter combinations&lt;/li&gt;
&lt;li&gt;View the logs to see the JSON-RPC messages being exchanged&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Use MCPJam Inspector's Advanced Features
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Chat Interface&lt;/strong&gt;: Interact with your MCP server conversationally using natural language&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LLM Playground&lt;/strong&gt;: Test how different LLMs (OpenAI, Claude, Ollama) use your MCP server's tools&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Resource Templates&lt;/strong&gt;: If available, explore prompt templates for standardized tool usage&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tracing&lt;/strong&gt;: Monitor detailed request/response flows to understand how the MCP protocol works&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test Cases&lt;/strong&gt;: Create and save test cases for automated testing of your MCP integrations&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Happy Coding!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>testing</category>
      <category>tooling</category>
    </item>
    <item>
      <title>A Smarter Way to Find and Test AI Models for Your App using GPT + Super AI (MCP)</title>
      <dc:creator>TheRabbitHole</dc:creator>
      <pubDate>Mon, 19 Jan 2026 16:26:09 +0000</pubDate>
      <link>https://dev.to/therabbithole/a-smarter-way-to-find-and-test-ai-models-for-your-app-using-gpt-super-ai-mcp-3639</link>
      <guid>https://dev.to/therabbithole/a-smarter-way-to-find-and-test-ai-models-for-your-app-using-gpt-super-ai-mcp-3639</guid>
      <description>&lt;p&gt;Modern development tools have made building applications easier than ever. You can now launch a new app with a database, authentication, and other core features in minutes. The final piece of the puzzle, adding genuine intelligence with AI, however, introduces a new set of challenges.&lt;/p&gt;

&lt;p&gt;Developers often face several key questions when integrating AI:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Which AI model provider should you choose?&lt;/li&gt;
&lt;li&gt;  How do you price your product to account for AI usage costs?&lt;/li&gt;
&lt;li&gt;  If you're using your own API key, how do you protect it from misuse and prevent unexpected expenses?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These questions become even more critical if you plan to offer a free trial or a free tier for your application. Without a proper strategy, you risk having your budget drained by overuse and users who don't intend to subscribe. While many solutions exist, one straightforward approach is to ship your product with a local AI that performs its specific task efficiently.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Power of Local, Specialized AI
&lt;/h3&gt;

&lt;p&gt;Amazing technologies are available that allow you to bundle a lightweight AI model directly with your application. This can be as simple as the snippet below, which creates a basic chatbot within a single HTML file.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdn10e8xkye32hmpj1b60.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fdn10e8xkye32hmpj1b60.png" width="800" height="313"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Before you adopt this approach, there are two fundamental questions you need to answer:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;What specific use case should your model excel at?&lt;/strong&gt; Most developers know that smaller models are not generalists like the mega-models behind services like ChatGPT. Instead, they are fast, cheap, and lightweight specialists. Your use case might be document classification, language translation, text summarization, or another focused task.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Which model is the right one for that use case?&lt;/strong&gt; After defining the task, you need to find a model that can perform it effectively.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The first requirement is a core part of any successful business plan. The second, however, can be a significant challenge when you have to choose from hundreds of available models. There are many benchmarking platforms like Hugging Face's LLM Leaderboard, LMSys's Chatbot Arena, and Artificial Analysis, plus countless online playgrounds to test individual models. But sifting through them all takes time.&lt;/p&gt;

&lt;h3&gt;
  
  
  Automating Model Discovery with AI
&lt;/h3&gt;

&lt;p&gt;If you have a handful of use cases and need to iterate quickly, you can use AI an &lt;a href="https://apify.com/flamboyant_leaf/super-ai-bench-mcp/api?ref=airabbit.blog" rel="noopener noreferrer"&gt;Super AI MCP&lt;/a&gt; to automate the discovery and testing process. Here’s how it works:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Configure an AI to access benchmark data.&lt;/strong&gt; This gives your AI assistant the information it needs to compare models.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Configure the AI to access prediction platforms.&lt;/strong&gt; This connects your AI to services that host a wide variety of models.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Provide your use cases in natural language.&lt;/strong&gt; Let the AI find the most suitable models and run tests for you.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbji93nqyk4cz2ngwklpt.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fbji93nqyk4cz2ngwklpt.jpeg" width="800" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;To make this work, you only need two key components:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Any chatbot that supports the &lt;strong&gt;Model Context Protocol (MCP)&lt;/strong&gt;, such as ChatGPT, Claude, and others.&lt;/li&gt;
&lt;li&gt;  A free account at &lt;strong&gt;Apify.com&lt;/strong&gt; to access benchmark data using a specific MCP. (Requires an API key).&lt;/li&gt;
&lt;li&gt;  (Optional) A &lt;strong&gt;Replicate&lt;/strong&gt; account if you want to run predictions. (Requires an API key).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can then use a prompt like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Find the best 3 small models that can do this task and try them out on Replicate: 

--- my task 1 here 
--- my task 2 here 
etc..
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Let’s walk through an example.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prerequisites:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  ChatGPT (or another chatbot with MCP support)&lt;/li&gt;
&lt;li&gt;  An Apify API key (a free account is sufficient)&lt;/li&gt;
&lt;li&gt;  A Replicate API key (this is a pay-per-use service)&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Step-by-Step Guide to Automated Model Testing
&lt;/h3&gt;

&lt;h4&gt;
  
  
  Step 1: Configure the MCP Server
&lt;/h4&gt;

&lt;p&gt;First, you need to connect your chatbot to the benchmark and prediction tools using an MCP server.&lt;/p&gt;

&lt;p&gt;Start by adding a new MCP in your chatbot's settings.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fc3mvg1uoequo6gpyxeg9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fc3mvg1uoequo6gpyxeg9.png" width="530" height="384"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You will need to provide the server URL.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzipqw059sb4ul7e372zv.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fzipqw059sb4ul7e372zv.jpeg" width="800" height="604"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Use the following URL, adding your Replicate API key at the end where indicated.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;https://flamboyant-leaf--super-ai-bench-mcp.apify.actor/mcp?replicateApiKey=&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F54dwuq9mn4v91jgmnb60.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F54dwuq9mn4v91jgmnb60.png" width="800" height="1214"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Leave the OAuth section empty, as you will authenticate with Apify later. Click confirm to save.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2g2phae5ux4a96byc5c8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2g2phae5ux4a96byc5c8.png" width="800" height="397"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That's it for the configuration.&lt;/p&gt;

&lt;h4&gt;
  
  
  Step 2: Find and Analyze Suitable Models
&lt;/h4&gt;

&lt;p&gt;Now, let's try a simple example to find some high-value small models. Later, you can replace this with your own specific use cases.&lt;/p&gt;

&lt;p&gt;In your chatbot, enter the following prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Find the best small model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;ChatGPT will now ask the benchmark tool for suitable models and sort them based on the request.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fumjid9qh1hspe74iligt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fumjid9qh1hspe74iligt.png" width="800" height="647"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here, it has found several models, including different versions of Llama, Qwen, and Phi, along with necessary data like size and cost.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fg4a4fpqy6n6a76j85det.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fg4a4fpqy6n6a76j85det.png" width="800" height="488"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The AI then provides a quick recommendation of which models to use.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Flf59wdqfk2j5m32i9ybe.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Flf59wdqfk2j5m32i9ybe.png" width="800" height="223"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Step 3: Test the Models on Replicate
&lt;/h4&gt;

&lt;p&gt;This is useful, but the real power comes from seeing the models execute your use case. Here, we'll let the AI create and run a simple coding task.&lt;/p&gt;

&lt;p&gt;Use the following prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;try them on replicate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The AI will first search for suitable models available on the Replicate platform. Note that not all models listed in benchmarks are on Replicate, but in this case, they are.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fquvd9pn7awu6blt2g4bq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fquvd9pn7awu6blt2g4bq.png" width="800" height="424"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Now, we can run the test on all of them simultaneously.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2ozek1cfdlilz9ekarlo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F2ozek1cfdlilz9ekarlo.png" width="800" height="435"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You can see the jobs running in your Replicate dashboard, with details including creation date, duration, and more. Your AI also has access to this data.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://replicate.com/predictions?ref=airabbit.blog" rel="noopener noreferrer"&gt;https://replicate.com/predictions&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fh3v86am0s0dtxytbnj3s.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fh3v86am0s0dtxytbnj3s.png" width="800" height="266"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After approximately one to two minutes, our use case has been tested across five different models, and we receive a detailed analysis directly from the AI.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fah1wauku9ogmbo7bgdto.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fah1wauku9ogmbo7bgdto.png" width="800" height="1333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Real-World Applications and Benefits
&lt;/h3&gt;

&lt;p&gt;This was a very simple example. In a real-world scenario, you can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  Provide your own complex, specific use cases for testing.&lt;/li&gt;
&lt;li&gt;  Save the results for future comparison.&lt;/li&gt;
&lt;li&gt;  Evaluate new models as they are released without switching between different platforms.&lt;/li&gt;
&lt;li&gt;  Distribute complex tasks across multiple models to leverage their unique strengths.&lt;/li&gt;
&lt;li&gt;  And much more.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;While all of these capabilities are valuable, the greatest benefit is the ability to quickly compare results from different models without subscribing to multiple services. As mentioned at the beginning of this post, this process makes it significantly easier to find small, efficient models that you can confidently ship with your products.&lt;/p&gt;

&lt;h2&gt;
  
  
  Appendix
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Pure HTML/JS Chatbot (Snippet)
&lt;/h3&gt;

&lt;p&gt;Open your Chrome browser and enable the on-device model at&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;chrome://flags/#optimization-guide-on-device-model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then save this HTML file and just open it. The rest is self-explanatory.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;&amp;lt;!doctype html&amp;gt;
&amp;lt;html lang="en"&amp;gt;
&amp;lt;head&amp;gt;
  &amp;lt;meta charset="utf-8" /&amp;gt;
  &amp;lt;meta name="viewport" content="width=device-width,initial-scale=1" /&amp;gt;
  &amp;lt;title&amp;gt;Local LLM Chat (Browser)&amp;lt;/title&amp;gt;
  &amp;lt;style&amp;gt;
    :root { color-scheme: dark; }
    body { margin: 0; font: 14px/1.4 system-ui, -apple-system, Segoe UI, Roboto, Arial; background:#0b0f14; color:#e6edf3; }
    .wrap { max-width: 980px; margin: 0 auto; padding: 16px; display:flex; flex-direction:column; gap:12px; height: 100vh; box-sizing:border-box; }
    .top { display:flex; gap:10px; align-items:center; flex-wrap:wrap; }
    .chip { padding:6px 10px; border:1px solid #223; border-radius:999px; background:#0f1621; }
    .status { opacity:.9; }
    .chat { flex:1; overflow:auto; border:1px solid #223; border-radius:12px; padding:12px; background:#0f1621; }
    .msg { margin: 0 0 10px 0; white-space:pre-wrap; }
    .msg .role { font-weight:700; }
    .msg.user .role { color:#7ee787; }
    .msg.ai .role { color:#79c0ff; }
    .row { display:flex; gap:10px; }
    input, select {
      padding:10px; border-radius:10px; border:1px solid #223;
      background:#0b0f14; color:#e6edf3;
    }
    #inp { flex:1; }
    button { padding:10px 12px; border-radius:10px; border:1px solid #223; background:#1f6feb; color:#fff; cursor:pointer; }
    button.secondary { background:#0f1621; }
    button:disabled { opacity:.5; cursor:not-allowed; }
    .small { font-size: 12px; opacity:.8; }
    .hide { display:none; }
  &amp;lt;/style&amp;gt;
&amp;lt;/head&amp;gt;
&amp;lt;body&amp;gt;
  &amp;lt;div class="wrap"&amp;gt;
    &amp;lt;div class="top"&amp;gt;
      &amp;lt;span class="chip"&amp;gt;Transformers.js (browser local)&amp;lt;/span&amp;gt;

      &amp;lt;label&amp;gt;
        Model:
        &amp;lt;select id="modelSelect"&amp;gt;
          &amp;lt;option value="HuggingFaceTB/SmolLM2-135M-Instruct"&amp;gt;SmolLM2-135M-Instruct (recommended)&amp;lt;/option&amp;gt;
          &amp;lt;option value="HuggingFaceTB/SmolLM2-360M-Instruct"&amp;gt;SmolLM2-360M-Instruct (bigger)&amp;lt;/option&amp;gt;
          &amp;lt;option value="HuggingFaceTB/SmolLM2-1.7B-Instruct"&amp;gt;SmolLM2-1.7B-Instruct (heavy)&amp;lt;/option&amp;gt;
          &amp;lt;option value="__custom__"&amp;gt;Custom model id…&amp;lt;/option&amp;gt;
        &amp;lt;/select&amp;gt;
      &amp;lt;/label&amp;gt;

      &amp;lt;input id="customModel" class="hide" placeholder="e.g. Org/RepoName" size="28" /&amp;gt;

      &amp;lt;button id="loadBtn" type="button"&amp;gt;Load&amp;lt;/button&amp;gt;
      &amp;lt;button id="clearBtn" type="button" class="secondary" disabled&amp;gt;Clear&amp;lt;/button&amp;gt;

      &amp;lt;span class="status" id="status"&amp;gt;Not loaded.&amp;lt;/span&amp;gt;
    &amp;lt;/div&amp;gt;

    &amp;lt;div class="chat" id="chat"&amp;gt;&amp;lt;/div&amp;gt;

    &amp;lt;div class="row"&amp;gt;
      &amp;lt;input id="inp" placeholder="Type a message and press Enter…" disabled /&amp;gt;
      &amp;lt;button id="sendBtn" type="button" disabled&amp;gt;Send&amp;lt;/button&amp;gt;
    &amp;lt;/div&amp;gt;

    &amp;lt;div class="small"&amp;gt;
      If opening as &amp;lt;code&amp;gt;file://&amp;lt;/code&amp;gt; blocks module imports on your machine, run a local server:
      &amp;lt;code&amp;gt;python -m http.server 8000&amp;lt;/code&amp;gt; then open &amp;lt;code&amp;gt;http://localhost:8000&amp;lt;/code&amp;gt;.
      First load downloads the model (can be large).
    &amp;lt;/div&amp;gt;
  &amp;lt;/div&amp;gt;

  &amp;lt;script type="module"&amp;gt;
    const $ = (id) =&amp;gt; document.getElementById(id);
    const chatEl = $("chat");
    const statusEl = $("status");
    const inp = $("inp");
    const sendBtn = $("sendBtn");
    const clearBtn = $("clearBtn");
    const loadBtn = $("loadBtn");
    const modelSelect = $("modelSelect");
    const customModel = $("customModel");

    function escapeHtml(s) {
      return String(s).replace(/[&amp;amp;&amp;lt;&amp;gt;"']/g, (c) =&amp;gt; ({
        "&amp;amp;":"&amp;amp;amp;","&amp;lt;":"&amp;amp;lt;","&amp;gt;":"&amp;amp;gt;",'"':"&amp;amp;quot;","'":"&amp;amp;#39;"
      }[c]));
    }

    function addMsg(role, text) {
      const div = document.createElement("div");
      div.className = `msg ${role}`;
      div.innerHTML = `&amp;lt;span class="role"&amp;gt;${role === "user" ? "You" : "AI"}:&amp;lt;/span&amp;gt; ${escapeHtml(text)}`;
      chatEl.appendChild(div);
      chatEl.scrollTop = chatEl.scrollHeight;
    }

    function setUiLoaded(loaded) {
      inp.disabled = !loaded;
      sendBtn.disabled = !loaded;
      clearBtn.disabled = !loaded;
    }

    modelSelect.addEventListener("change", () =&amp;gt; {
      const isCustom = modelSelect.value === "__custom__";
      customModel.classList.toggle("hide", !isCustom);
    });

    // Chat state
    let generator = null;
    let deviceUsed = "";
    const system = "System: You are a helpful assistant. Be concise.\n";
    let transcript = "";

    function resetChat() {
      transcript = "";
      chatEl.innerHTML = "";
      addMsg("ai", "Ready. Ask me a question.");
      inp.focus();
    }

    async function loadModel() {
      try {
        setUiLoaded(false);
        loadBtn.disabled = true;
        statusEl.textContent = "Loading library…";

        const { pipeline, env } = await import(
          "https://cdn.jsdelivr.net/npm/@huggingface/transformers@3.0.2/+esm"
        );
        env.useBrowserCache = true;

        let modelId = modelSelect.value;
        if (modelId === "__custom__") modelId = customModel.value.trim();
        if (!modelId) throw new Error("No model id provided.");

        const make = async (device) =&amp;gt; pipeline("text-generation", modelId, {
          dtype: "q4",
          device,
          progress_callback: (p) =&amp;gt; {
            if (p &amp;amp;&amp;amp; p.status === "progress") {
              const pct = (typeof p.progress === "number") ? ` ${p.progress.toFixed(1)}%` : "";
              statusEl.textContent = `Downloading ${p.file || ""}${pct}`.trim();
            }
          },
        });

        try {
          statusEl.textContent = "Initializing WebGPU…";
          generator = await make("webgpu");
          deviceUsed = "webgpu";
        } catch (e) {
          statusEl.textContent = "WebGPU failed, using WASM…";
          generator = await make("wasm");
          deviceUsed = "wasm";
        }

        statusEl.textContent = `Loaded ${modelId} (${deviceUsed}).`;
        setUiLoaded(true);
        resetChat();
      } catch (e) {
        console.error(e);
        statusEl.textContent = `Load failed: ${e.message || e}`;
        addMsg("ai", "Load failed. Check console. If using file:// and imports are blocked, run via a local server.");
        generator = null;
        deviceUsed = "";
        setUiLoaded(false);
      } finally {
        loadBtn.disabled = false;
      }
    }

    async function send() {
      if (!generator) return;

      const user = inp.value.trim();
      if (!user) return;

      inp.value = "";
      addMsg("user", user);

      transcript += `User: ${user}\nAssistant:`;
      statusEl.textContent = "Thinking…";
      sendBtn.disabled = true;
      inp.disabled = true;

      try {
        const out = await generator(system + transcript, {
          max_new_tokens: 160,
          temperature: 0.7,
          return_full_text: false
        });

        const r = Array.isArray(out) ? out[0] : out;
        const aiText = (r &amp;amp;&amp;amp; r.generated_text != null) ? String(r.generated_text).trim() : "";
        transcript += ` ${aiText}\n`;

        addMsg("ai", aiText || "(no output)");
        statusEl.textContent = `Loaded (${deviceUsed}).`;
      } catch (e) {
        console.error(e);
        statusEl.textContent = "Generation error (see console).";
        addMsg("ai", "Error generating response. See console.");
      } finally {
        sendBtn.disabled = false;
        inp.disabled = false;
        inp.focus();
      }
    }

    sendBtn.addEventListener("click", send);
    inp.addEventListener("keydown", (e) =&amp;gt; { if (e.key === "Enter") send(); });
    clearBtn.addEventListener("click", resetChat);
    loadBtn.addEventListener("click", loadModel);

    // Optional: auto-load on open
    // loadModel();
  &amp;lt;/script&amp;gt;
&amp;lt;/body&amp;gt;
&amp;lt;/html&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
    </item>
    <item>
      <title>The End of AI Monogamy: Let AI Find the Best Model for Your Task</title>
      <dc:creator>TheRabbitHole</dc:creator>
      <pubDate>Fri, 16 Jan 2026 11:24:23 +0000</pubDate>
      <link>https://dev.to/therabbithole/the-end-of-ai-monogamy-let-ai-find-the-best-model-for-your-task-23p7</link>
      <guid>https://dev.to/therabbithole/the-end-of-ai-monogamy-let-ai-find-the-best-model-for-your-task-23p7</guid>
      <description>&lt;p&gt;Most of us spend an insane amount of time using AI. Whether it's coding, writing, or analyzing data, we are glued to our prompts. But here is the problem: &lt;strong&gt;We are almost all "monogamous" with our AI.&lt;/strong&gt; You probably have a subscription to ChatGPT, or maybe Claude, or Gemini. You know deep down that other models exist. You know that for certain tasks, a specialized model like DeepSeek or Llama 3 might be faster, cheaper, or smarter. But you don't switch. &lt;/p&gt;

&lt;p&gt;Why? &lt;br&gt;
Maybe it's not just the hassle of jumping into a new playground. &lt;br&gt;
Or maybe It's that &lt;strong&gt;generic benchmarks rarely match reality.&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;We see leaderboards claiming a model is "#1 in Coding," but that is based on a standardized dataset. It doesn't tell you if the model is good at &lt;em&gt;your&lt;/em&gt; specific legacy code, your unique tone of voice, or your particular data structure. A global average is meaningless when you have a specific problem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This is inefficient.&lt;/strong&gt; Relying on a general-purpose winner for every single task is a compromise. What if you didn't have to guess? What if your current AI assistant could run a "micro-benchmark" for you—using your actual prompt—right in the middle of your conversation?&lt;/p&gt;

&lt;h2&gt;
  
  
  The "Auto-Pilot" Benchmark
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Example 1: Legacy Code Refactoring (Python)&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;You have a 500-line Django ORM query that's killing your database performance. Instead of asking ChatGPT and hoping:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"I have a 500-line Django ORM query that's killing our database performance. Run this code snippet through the top 3 LLM models on Replicate and show me their refactoring approaches side-by-side."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Why this works:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude might suggest async queries&lt;/li&gt;
&lt;li&gt;DeepSeek might catch a specific database indexing issue&lt;/li&gt;
&lt;li&gt;Llama might propose a completely different query structure&lt;/li&gt;
&lt;li&gt;You see all three perspectives &lt;strong&gt;in parallel&lt;/strong&gt; instead of re-prompting 3 times&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  &lt;strong&gt;Example 2: Data Analysis on Your Real Dataset&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;You have actual sales data and need insights:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Here's my Q4 sales CSV. Find the top 3 models best at statistical reasoning, send them this data, and show me which model catches the most actionable insights."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Why this works:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GPT-4 might focus on trend analysis&lt;/li&gt;
&lt;li&gt;Claude might catch subtle correlations you missed&lt;/li&gt;
&lt;li&gt;Llama might be faster/cheaper and still identify key patterns&lt;/li&gt;
&lt;li&gt;You're benchmarking on &lt;strong&gt;YOUR data&lt;/strong&gt;, not generic datasets&lt;/li&gt;
&lt;/ul&gt;




&lt;h3&gt;
  
  
  &lt;strong&gt;Example 3: Multilingual Content with Brand Voice Matching&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;You need marketing copy in multiple languages with a specific tone:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Write marketing copy for our premium SaaS in English, German, and Japanese. First, query which models are best at multilingual tone-matching, then run the same prompt through the top 2 models and show me the differences."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;Why this works:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You see if one model nails your brand voice better&lt;/li&gt;
&lt;li&gt;Some models are objectively better at specific languages&lt;/li&gt;
&lt;li&gt;You pick the winner for each language instead of settling for one model&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  How It Works Under the Hood
&lt;/h2&gt;

&lt;p&gt;By connecting an &lt;a href="https://console.apify.com/actors/1hdn3N9PtIi5z4ePY/information/latest/readme" rel="noopener noreferrer"&gt;MCP&lt;/a&gt; (Model Context Protocol) client to live data sources, we bridge the gap between static leaderboards and active workflows.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Context Awareness:&lt;/strong&gt; The AI detects if you are doing creative writing, logic puzzles, or hardcore engineering.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Lookup:&lt;/strong&gt; It queries the benchmark tool to find the highest-performing models for that specific category.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Execution:&lt;/strong&gt; It uses the Replicate API to spin up instances of those top models, feeds them your prompt, and aggregates the results.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;You get 3 or 4 distinct answers from the smartest models on the planet, tailored exactly to the problem you are solving right now.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Disclaimer&lt;/strong&gt;: The tools and workflows presented in this article provide a preliminary glimpse into the performance of various AI models, but these results should not be taken for granted. Automated comparisons are illustrative and may not reflect performance across all scenarios. To fully understand the specific strengths and weaknesses of candidate models, you must independently verify the results against your own data and requirements.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Getting Started
&lt;/h2&gt;

&lt;p&gt;You only need two things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Apify Account:&lt;/strong&gt; Powers the benchmark scraping. Free account gives you &lt;strong&gt;$5/month in credits&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Replicate Account:&lt;/strong&gt; Provides access to models. &lt;strong&gt;Pay-per-use&lt;/strong&gt;, no monthly fees.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Step 1: Configure Your MCP Client
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ai-live-benchmark"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"npx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"-y"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"mcp-remote"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"https://flamboyant-leaf--super-ai-bench-mcp.apify.actor/mcp"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"--header"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"Authorization: Bearer &amp;lt;APIFY_API_TOKEN&amp;gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"--header"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"X-Replicate-API-Key: &amp;lt;REPLICATE_API_KEY&amp;gt;"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Step 2: Run the Test&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Now for the fun part. We don't need to specify which benchmark to use. We just give the AI a task. Let’s try a specific multilingual request:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;" Before you start, Read the Documentation. Then Find the 3 most powerful LLM models and run on Replicate to do this task: Write an email to my boss excusing being late in German."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Here is what happens next in real-time:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase A: The Smart Lookup&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;First, the AI analyzes your request. It realizes this is a text generation task involving a foreign language. It automatically decides to query the benchmark API for the current top-performing Large Language Models.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7qn74esv5lios8gy770o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2F7qn74esv5lios8gy770o.png" width="800" height="341"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase B: Finding the Models&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Next, it takes those top-ranked models and searches the Replicate "Model Garden" to see which ones are available for immediate access.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(Note: Sometimes a specific model version might not be hosted on Replicate. In that case, the agent is smart enough to just pick the next best model from the benchmark list—or you can simply ask it to "try the next one.")&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmg76owqg0iqep76h9wsq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmg76owqg0iqep76h9wsq.png" width="800" height="459"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase C: The Live Showdown&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Finally, it runs the prediction. It doesn't just give you one answer; it executes the task on all three models in parallel.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fy9wc2lkwka7dpxwsb9rf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fy9wc2lkwka7dpxwsb9rf.png" width="800" height="374"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmaifn94os20xiiadphfn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fmaifn94os20xiiadphfn.png" width="800" height="494"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Flczbj3kp0h8qa9bcddqf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Flczbj3kp0h8qa9bcddqf.png" width="800" height="296"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Please note that sometimes Claude or another AI might guess the 'best' model by itself and start searching for it on Repclaiase. To avoid this, tell it explicitly to look up suitable benchmarks and let it search without outputting the result. This will give you a better understanding of what it is doing under the hood and what is suitable for your specific use cases.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Final thought
&lt;/h3&gt;

&lt;p&gt;This isn't limited to emails or code. This workflow fully supports &lt;strong&gt;Image Models&lt;/strong&gt; (Nano Banana, Qwen Image etc.) too. You can ask it to "Generate a cyberpunk city using the top 3 image models," and you will get a side-by-side comparison of Flux, Stable Diffusion, and others in one shot. And if you are using an interface like &lt;strong&gt;Claude Artifacts&lt;/strong&gt; or &lt;strong&gt;Canvas&lt;/strong&gt;, you can even ask the AI to build a simple HTML gallery to display these results side-by-side for a true "blind taste test." But that’s a topic for another post!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
