<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Steven Gonsalvez</title>
    <description>The latest articles on DEV Community by Steven Gonsalvez (@stevengonsalvez).</description>
    <link>https://dev.to/stevengonsalvez</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F613704%2F6fa5896f-e719-469e-812f-d2f8f9971b0e.png</url>
      <title>DEV Community: Steven Gonsalvez</title>
      <link>https://dev.to/stevengonsalvez</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/stevengonsalvez"/>
    <language>en</language>
    <item>
      <title>Kimi K3 vs Fable vs Sol on a Normal Debug: The Cheapest Model Wasn't Cheap</title>
      <dc:creator>Steven Gonsalvez</dc:creator>
      <pubDate>Thu, 23 Jul 2026 21:54:06 +0000</pubDate>
      <link>https://dev.to/stevengonsalvez/kimi-k3-vs-fable-vs-sol-on-a-normal-debug-the-cheapest-model-wasnt-cheap-4bki</link>
      <guid>https://dev.to/stevengonsalvez/kimi-k3-vs-fable-vs-sol-on-a-normal-debug-the-cheapest-model-wasnt-cheap-4bki</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz15k3y7o69f8gfqso2kq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz15k3y7o69f8gfqso2kq.png" alt="The whole story in one sketchnote: a Mac scare that turned out to be Apple Continuity, three models reaching the same answer, and the bill showing the cheapest-per-token model was the dearest to run" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;TL;DR&lt;/p&gt;

&lt;p&gt;Same real debug, three models, one answer. The cheapest model per token (Kimi K3) cost about the same to actually run as the pricey two, and took more than twice as long. Price your models by the outcome, not the token.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Tokens&lt;/th&gt;
&lt;th&gt;Fresh / Cached / Output&lt;/th&gt;
&lt;th&gt;Est. cost&lt;/th&gt;
&lt;th&gt;Wall-clock&lt;/th&gt;
&lt;/tr&gt;&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;~800K&lt;/td&gt;
&lt;td&gt;60K / 540K / 200K&lt;/td&gt;
&lt;td&gt;~$6.60&lt;/td&gt;
&lt;td&gt;~21 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K3&lt;/td&gt;
&lt;td&gt;~2.9M&lt;/td&gt;
&lt;td&gt;240K / 2.185M / 475K&lt;/td&gt;
&lt;td&gt;~$8.50&lt;/td&gt;
&lt;td&gt;~51 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5&lt;/td&gt;
&lt;td&gt;~650K&lt;/td&gt;
&lt;td&gt;49K / 439K / 163K&lt;/td&gt;
&lt;td&gt;~$9.10&lt;/td&gt;
&lt;td&gt;~27 min&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Cheapest per token, dearest to run. Token totals and times are real; the split and cache-hit rate are approximations.&lt;/p&gt;




&lt;p&gt;Right, this one isn't a shotclubhouse ops story. This was my own Mac, my own Saturday.&lt;/p&gt;

&lt;p&gt;Here is what kept happening. Every fifteen minutes or so, &lt;a href="https://arc.net" rel="noopener noreferrer"&gt;Arc&lt;/a&gt; would throw up a &lt;a href="https://tailscale.com" rel="noopener noreferrer"&gt;Tailscale&lt;/a&gt; login prompt. On its own. Nobody touched anything. It never actually logged in, it just appeared, sat there, and buggered off. And that same Arc window had the little "your screen is being shared" dot lit up in the menu bar. A login prompt I didn't ask for, plus a sharing indicator I couldn't explain, on a clock. Not a great look.&lt;/p&gt;

&lt;p&gt;A bit scary. Could be a compromise. So the first thing on my mind wasn't "debug it", it was how do I stop the bleeding, without torching the evidence. No panic-reboot, no uninstalling Arc or Tailscale, because that wipes the exact volatile state you need to work out what happened. Contain, don't cremate.&lt;/p&gt;

&lt;p&gt;My first go-to was &lt;a href="https://www.anthropic.com" rel="noopener noreferrer"&gt;Claude Fable 5&lt;/a&gt;, on xhigh. I wrote the problem up as one cold prompt, no hints, no reveal, and I didn't want a lecture on what I &lt;em&gt;could&lt;/em&gt; check. I wanted it to actually do the work. So I turned it loose in an agentic harness and let it run the whole triage itself, contain, dig, trace, find the cause, no hand-holding from me. Fable drove through it end to end, landed on the cause, and eased my concern enough that I could go and reproduce the situation myself. Once I had it reproducing on demand, I turned &lt;a href="https://openai.com" rel="noopener noreferrer"&gt;GPT-5.6 "Sol"&lt;/a&gt; and &lt;a href="https://www.moonshot.ai" rel="noopener noreferrer"&gt;Kimi K3&lt;/a&gt; loose on the exact same cold prompt, same harness, same free rein, the cheap model against the two pricey ones.&lt;/p&gt;

&lt;p&gt;Why put K3 in the ring at all? Because it just landed, and it's a bit of a moment. Kimi K3 is open-weight, and on the &lt;a href="https://artificialanalysis.ai" rel="noopener noreferrer"&gt;Artificial Analysis Intelligence Index&lt;/a&gt; it sits third, at 57, behind only Claude Fable 5 and GPT-5.6 Sol, ahead of Claude Opus 4.8 and most of the GPT-5.6 line. An open model, top three, at a third of the price of the closed ones next to it. On paper that's a bit ridiculous.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh94ijevyw345z4h8gjel.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh94ijevyw345z4h8gjel.png" alt="Kimi K3 sits third on the Artificial Analysis Intelligence Index at 57, the top open-weight model, behind only Claude Fable 5 and GPT-5.6 Sol" width="799" height="449"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So the question was never whether it's smart. The benchmark says it is. The question is what that intelligence costs to cash in on a real job, not a puzzle, not a leaderboard eval, an actual "is my machine compromised" panic with a clock running.&lt;/p&gt;

&lt;p&gt;I'll say this up front so you can enjoy it on the way through: the root cause turned out to be really silly, and on hindsight I end up looking like a bit of an idiot. Worth it.&lt;/p&gt;

&lt;p&gt;The prompt, roughly: &lt;em&gt;Mac, possible compromise, Arc keeps opening a Tailscale login every 15 minutes, screen-sharing indicator is on. Contain it safely first, then investigate and find what is triggering it. Do the work yourself, don't just hand me the fix.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How Sol and Fable operated
&lt;/h2&gt;

&lt;p&gt;Both of them, credit where due, refused to jump to the fix, and neither sat there handing me a checklist to go run. They just did it. Wired into the Mac through &lt;a href="https://github.com/steipete/peekaboo" rel="noopener noreferrer"&gt;Peekaboo&lt;/a&gt; for screenshots and screen reads, plus &lt;a href="https://www.hammerspoon.org" rel="noopener noreferrer"&gt;Hammerspoon&lt;/a&gt; and a bit of macOS-control tooling for the actual clicking, each one went straight for the sharing dot itself, because macOS names the exact app capturing the screen the moment that menu opens, and that one move collapses half the mystery. From there both went hunting for a 15-minute launcher (&lt;code&gt;StartInterval&lt;/code&gt; of 900 seconds is the tell), and both worked to rule out the real remote-access gunk, VNC on port 5900, Screen Sharing, Remote Management, an MDM profile quietly bossing the machine about. Screenshots, &lt;code&gt;log show&lt;/code&gt;, &lt;code&gt;launchctl&lt;/code&gt;, the lot, driven by the model, not by me.&lt;/p&gt;

&lt;p&gt;(On hindsight I should have just pointed the &lt;a href="https://openai.com/codex" rel="noopener noreferrer"&gt;Codex app&lt;/a&gt;'s computer-use at the whole thing, apparently that's the top dog for driving a desktop right now. But I live in a terminal, I barely open an IDE these days, so it was agentic harness plus CLI tools or nothing.)&lt;/p&gt;

&lt;p&gt;But the interesting bit is where they &lt;em&gt;diverged&lt;/em&gt;, because it's a proper forensics fork.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sol went isolate-first.&lt;/strong&gt; Except it could hardly yank the network, that was the wire it was talking to me over. So the containment was surgical instead of blunt: clamp the firewall to a single allowlist and drop everything else. Default-deny egress with &lt;code&gt;pfctl&lt;/code&gt;, one hole punched for &lt;code&gt;api.openai.com&lt;/code&gt; on 443 plus DNS, and nothing else leaves the box. A live attacker or a remote-control channel now has nowhere to send to, while the model keeps the one uplink it needs to stay alive and keep working. Kill Bluetooth, stop typing anything sensitive, lock the screen but keep the Mac awake, then snapshot the host with that single pipe still up. Very "assume the attacker is live, close every road out except mine."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fable went snapshot-first.&lt;/strong&gt; And this is the sharper instinct, honestly. Fable's whole point: the moment you clamp that firewall shut, you lose the evidence of &lt;em&gt;where the data was going&lt;/em&gt;. So it grabbed the live sockets &lt;em&gt;before&lt;/em&gt; narrowing the egress down to its own uplink:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sudo lsof -i -P -n &amp;gt; ~/Desktop/ir-lsof.txt
netstat -anv    &amp;gt; ~/Desktop/ir-netstat.txt
ps auxww        &amp;gt; ~/Desktop/ir-ps.txt
launchctl list  &amp;gt; ~/Desktop/ir-launchctl.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every process with an open socket and its remote endpoint. If your screen is actually leaving the box, the sender is sitting right there as an established connection with sustained traffic. &lt;em&gt;Then&lt;/em&gt; it locked egress down to its own endpoint and nothing else. It is the difference between "stop the intruder" and "photograph the intruder leaving before you lock the door", and for a real incident, Fable's ordering is the one I would teach.&lt;/p&gt;

&lt;p&gt;Both then converged on the same decision tree: is this the benign end (a legit Tailscale client with a dead session politely re-opening its login page every retry, plus some app holding a screen-recording permission it isn't really using) or the malicious end (an unknown launchd job plus a remote-access tool)? And both were clear: a screen indicator on its own, a Tailscale permission on its own, a fixed 15-minute cadence on its own, none of those are proof of anything. You need to catch one popup in the act with the initiating process attached. Nice discipline. Suspicious until explained.&lt;/p&gt;

&lt;h2&gt;
  
  
  Kimi K3 on the stand
&lt;/h2&gt;

&lt;p&gt;Same cold prompt, &lt;a href="https://www.moonshot.ai" rel="noopener noreferrer"&gt;Kimi K3&lt;/a&gt;, same harness. And to be fair to it, the run is &lt;em&gt;good&lt;/em&gt;. Clean, well-structured, easily the most readable of the three. Five phases, contain, snapshot, trace the Tailscale prompt, trace the screen-share, find the 15-minute timer, and a proper "what this tells us" note under every single command. If you handed this to someone who had never triaged anything, they could follow it. It hit all the same beats: &lt;code&gt;lsof -i&lt;/code&gt; for the sockets, Arc history and &lt;code&gt;log show&lt;/code&gt; for &lt;code&gt;login.tailscale.com&lt;/code&gt; to find the parent process that opened the window, Screen Recording permissions plus WindowServer logs for the sharing dot, &lt;code&gt;launchd&lt;/code&gt; and a &lt;code&gt;StartInterval&lt;/code&gt; of 900 for the timer. It even hedged honestly in one spot ("if this yields nothing, we may need to adjust the predicate"). Solid.&lt;/p&gt;

&lt;p&gt;But two things separate a good-looking run from a good incident response.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It fumbled the one ordering that matters.&lt;/strong&gt; K3 led with clamping the network shut &lt;em&gt;first&lt;/em&gt;, and treated snapshotting the live connections as an optional afterthought, gated, bizarrely, on whether I had "an external drive or USB stick handy." That is exactly backwards, and the USB bit is nonsense, you save the output to the local disk. This is the precise step Fable built its whole run around: the moment you clamp the firewall shut, you lose the evidence of &lt;em&gt;where the data was going&lt;/em&gt;. K3 knew the step existed and then buried it under a caveat that would, in a real incident, cost you the single most important artefact. Cheap sticker. Not a cheap mistake.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No adversarial layer.&lt;/strong&gt; K3 ran a straight linear checklist. Sol brought a red team and a blue team that argued with each other and named their own proof gap; Fable brought an explicit benign-vs-malicious decision tree and an evidence threshold ("a single indicator isn't proof"). K3 gathers the evidence but never builds in the bit that stops it believing its own scary story. For a tidy report, fine. For a real "am I compromised" moment, that self-doubt is the feature you actually need.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually ran
&lt;/h2&gt;

&lt;p&gt;All three ran autonomously, but Sol is the one I've the full receipts for, because I drove it through &lt;a href="https://openai.com/codex" rel="noopener noreferrer"&gt;Codex&lt;/a&gt; and it left a transcript. Sol span up a two-agent team, a &lt;code&gt;red_team_triage&lt;/code&gt; cell to hunt for compromise and a &lt;code&gt;blue_team_validation&lt;/code&gt; cell to argue with it, and let them go at the machine.&lt;/p&gt;

&lt;p&gt;The verdict it came back with, and I'm quoting the shape of it: "automation/configuration collision, not compromise. Likelihood low, confidence 85%." The evidence it actually turned up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Two Tailscales fighting. A Homebrew &lt;code&gt;tailscaled&lt;/code&gt; on 1.86.2 &lt;em&gt;and&lt;/em&gt; the signed GUI plus system extension on 1.94.1, both live at once. Duplicate control paths, version skew, the single most likely source of the confusing behaviour. It even caught the gotcha that the Tailscale app binary launches the GUI instead of the CLI unless you force &lt;code&gt;TAILSCALE_BE_CLI=1&lt;/code&gt;, which is exactly the kind of thing that spawns a phantom window.&lt;/li&gt;
&lt;li&gt;SIP was disabled. Not evidence of compromise, but a fat exposure that makes everything worse if something &lt;em&gt;did&lt;/em&gt; run.&lt;/li&gt;
&lt;li&gt;Eight Arc opens between 10:05 and 10:09 with no referrer. Local Node and LaunchServices activity lined up in time, but, and this is the honest part, it &lt;em&gt;couldn't&lt;/em&gt; pin the initiating PID. Correlation, not a culprit.&lt;/li&gt;
&lt;li&gt;Everything reassuring otherwise: Arc and Tailscale properly signed and notarised, tailnet healthy, no exit node, no dodgy persistence, no MDM, the high-permission extensions all disabled.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And then the thing I actually respect most: the blue-team cell openly disagreed with the red-team cell. "Temporal correlation doesn't prove Node opened those URLs. No referrer doesn't prove malicious." It named its own proof gap out loud, capture one opening with the PID, executable path, parent process and mechanism, or you've got nothing. It refused to hand me a satisfying story it couldn't back. That is the good stuff. That is what you want a security agent to do instead of confidently screaming "COMPROMISED" at a purple dot.&lt;/p&gt;

&lt;h2&gt;
  
  
  So I caught it in the act
&lt;/h2&gt;

&lt;p&gt;Every one of those plans ended on the same note: stop theorising, catch one popup live with proof attached. So that's exactly what I did. I sat and waited for the fifteen-minute tick with the mouse plugged back in, and this time I actually looked.&lt;/p&gt;

&lt;p&gt;First, Control Centre:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftxjghw0kirzgdqfgndyi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftxjghw0kirzgdqfgndyi.png" alt="Control Centre showing Universal Control linking keyboard and mouse to Steven's MacBook Pro" width="320" height="325"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Link keyboard and mouse to: Steven's MacBook Pro.&lt;/em&gt; &lt;a href="https://support.apple.com/en-us/102459" rel="noopener noreferrer"&gt;Universal Control&lt;/a&gt;, live. My keyboard and mouse weren't just mine, they were bridged straight to the MacBook Pro sitting three feet away.&lt;/p&gt;

&lt;p&gt;Then the Arc window popped on the timer, right on cue. And this time I looked at the dock icon properly instead of the scary prompt:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2r7yk4mn3na7ewu20ilt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2r7yk4mn3na7ewu20ilt.png" alt="Arc dock icon badged with a laptop, labelled Arc from MacBook Pro" width="186" height="213"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A little laptop badge. &lt;em&gt;Arc, From MacBook Pro.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;There it's, in one tooltip. That badge isn't screen-sharing, it's &lt;a href="https://support.apple.com/en-us/HT204681" rel="noopener noreferrer"&gt;Apple Continuity&lt;/a&gt;, the Mac-answers-your-iPhone family of features, advertising the MacBook Pro's Arc session on this screen. Arc on the other Mac had a Tailscale login open, Continuity handed the session across, and because my keyboard and mouse were bridged over through Universal Control, a stray focus change popped it onto this screen. Tailscale was just the page being carried, never a compromise. Cut the cross-Mac link and the popups stop. Two Macs in the same room, the second one basically my headless server that I had forgotten still had a screen on it and Continuity quietly switched on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reckoning: what it actually cost
&lt;/h2&gt;

&lt;p&gt;Here is the part that reframes the whole thing. All three identified the same cause, a benign Apple Continuity collision between the two Macs, not a breach. But look at what each one &lt;em&gt;spent&lt;/em&gt; getting there.&lt;/p&gt;

&lt;p&gt;First, the rate cards, because they aren't the same shape. Every vendor gives you a big discount on cached input, but the cached rate itself differs, and so does the output rate, which is where a lean run spends most of its money:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input / 1M&lt;/th&gt;
&lt;th&gt;Cached input / 1M&lt;/th&gt;
&lt;th&gt;Output / 1M&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;$30.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;$1.00&lt;/td&gt;
&lt;td&gt;$50.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K3&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$3.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.30&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$15.00&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Now the runs. A triage like this is a tool-heavy agent loop, it re-feeds a big stable prefix (system prompt, the growing transcript, tool output) on every single turn, so most &lt;em&gt;input&lt;/em&gt; is a cache hit. Sol and Fable ran lean, call it a 75/25 input-output split. Kimi K3 was a different shape entirely, which is the whole point in a second. Prices are each vendor's &lt;a href="https://www.anthropic.com/pricing" rel="noopener noreferrer"&gt;published July 2026 rates&lt;/a&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Tokens&lt;/th&gt;
&lt;th&gt;Fresh in / Cached in / Output&lt;/th&gt;
&lt;th&gt;Est. cost*&lt;/th&gt;
&lt;th&gt;Wall-clock&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;~800K&lt;/td&gt;
&lt;td&gt;60K / 540K / 200K&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$6.60&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~21 min&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K3&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~2.9M&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;240K / 2.185M / 475K&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$8.50&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~51 min&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5&lt;/td&gt;
&lt;td&gt;~650K&lt;/td&gt;
&lt;td&gt;49K / 439K / 163K&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$9.10&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;~27 min&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;*Illustrative, not a receipt. Token totals and times are real; the split and cache-hit rate are approximations. K3 ran through Claude Code, not Moonshot's own harness, with no API errors, so the token pile and the clock are the model's, not a broken pipe.&lt;/p&gt;

&lt;p&gt;So the cheapest model on the card, K3, cost about $8.50. Fable, the priciest per-token by a mile, cost $9.10. Sol landed at $6.60. Three wildly different stickers, three bills in the same narrow band.&lt;/p&gt;

&lt;p&gt;That is the point. K3 burned 2.9 million tokens and fifty-one minutes to reach the verdict Sol got in 800K and twenty-one. Its cheap rate didn't make the job cheap, it hid the cost in cached input and in wall-clock. By tokens, K3 wins. By outcome, the thing you actually paid for, all three cost about the same, and K3 was the slowest by a mile. Do not price a model by its tokens. Price it by what it takes to get you the answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The verdict
&lt;/h2&gt;

&lt;p&gt;On method: Fable had the sharpest instinct, snapshot the live state &lt;em&gt;before&lt;/em&gt; you clamp the firewall, because you can't un-destroy evidence. Sol ran the most thorough of the three, a red team and a blue team, and had the spine to flag what it had not proven. K3 wrote the most readable write-up and took the slowest, priciest road to the same place. All three got it right, and all three landed on the cause, a second Mac handing its window across through Apple Continuity. I reproduced it myself just to nail the exact toggle.&lt;/p&gt;

&lt;p&gt;And the lesson underneath it: the cheapest model on paper wasn't the cheapest to run. &lt;strong&gt;And price your models by the outcome, not the token.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>tailscale</category>
      <category>macos</category>
      <category>incidentresponse</category>
      <category>gpt56</category>
    </item>
    <item>
      <title>Anthropic Just Admitted MCP Has a Context Problem</title>
      <dc:creator>Steven Gonsalvez</dc:creator>
      <pubDate>Fri, 10 Jul 2026 15:44:43 +0000</pubDate>
      <link>https://dev.to/stevengonsalvez/anthropic-just-admitted-mcp-has-a-context-problem-1ona</link>
      <guid>https://dev.to/stevengonsalvez/anthropic-just-admitted-mcp-has-a-context-problem-1ona</guid>
      <description>&lt;h2&gt;
  
  
  Anthropic Built a Fix That Proves the Problem 🔍
&lt;/h2&gt;

&lt;p&gt;Right, so Anthropic dropped Tool Search on November 24th alongside Claude Opus 4.5, and I need you to sit with the implications for a second because they're brutal for MCP.&lt;/p&gt;

&lt;p&gt;Tool Search lets you mark tools as &lt;code&gt;defer_loading: true&lt;/code&gt;. When you do that, the tool's schema never enters your context window until Claude actually needs it. Claude gets a name and nothing else. When a task comes in that requires the tool, Claude calls the Tool Search tool, pulls the schema, and only &lt;em&gt;then&lt;/em&gt; does it have the full definition loaded. Lazy loading for AI tools. Sounds boring. It's not.&lt;/p&gt;

&lt;p&gt;Here's the number that should make you spit your tea out: accuracy went from 49% to 74% on Opus 4. On Opus 4.5, it climbed from 79.5% to 88.1%. Same model. Same tools. The only difference is &lt;em&gt;not loading the tool definitions upfront&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;I'll say it again. The mere presence of tool schemas in context was tanking accuracy by 25 percentage points. Twenty-five. Not because the model got confused by irrelevant tools (though it did). Because the context window was full of JSON nobody was using, and the model was drowning in it.&lt;/p&gt;




&lt;h2&gt;
  
  
  This is an admission dressed up as a feature
&lt;/h2&gt;

&lt;p&gt;I've been banging on about &lt;a href="https://dev.to/byte-sized-banter/week-28-mcp-context-kryptonite"&gt;MCP's context cost&lt;/a&gt; since July. The Playwright MCP server alone burns 12.8k tokens just sitting there. Load four or five MCP servers and you've torched half your context window before asking the model to do anything useful.&lt;/p&gt;

&lt;p&gt;Tool Search is Anthropic's answer. And the answer is: don't load the tools.&lt;/p&gt;

&lt;p&gt;Think about that for a second. The company that &lt;em&gt;created&lt;/em&gt; MCP built a feature whose entire purpose is avoiding the cost of loading MCP tool definitions. If that isn't an admission that the protocol has a context problem, I don't know what is. They didn't say "we've made MCP more efficient." They said "here's a way to not load MCP schemas until the last possible moment." The fix &lt;em&gt;is&lt;/em&gt; the diagnosis.&lt;/p&gt;

&lt;p&gt;And the timing is proper interesting. Three weeks earlier, on November 5th, mcporter dropped. mcporter converts MCP servers into plain code. Regular functions. No protocol overhead, no JSON schemas burning context, just code your agent calls directly. Two completely different approaches to the same problem, landing in the same month. The ecosystem is routing around MCP overhead from both directions simultaneously: Anthropic from the top (defer the loading) and the community from the bottom (skip the protocol entirely).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;📚 Geek Corner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Deferred loading vs. elimination&lt;/strong&gt;: Tool Search and mcporter solve the same problem differently. Tool Search keeps MCP but hides it until needed (lazy loading). mcporter removes MCP entirely and converts tool definitions to native code (compilation). The lazy loading approach preserves the MCP ecosystem but still pays the schema cost when tools activate. The compilation approach pays zero ongoing cost but loses MCP's dynamic discovery. Both are concessions that the original "load everything upfront" model was broken.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  The real question nobody's asking
&lt;/h2&gt;

&lt;p&gt;If deferred loading improves accuracy by 25 percentage points, what does that tell us about every benchmark result published before November 24th? All those MCP-heavy setups were being evaluated with bloated context. Accuracy scores were lower than they needed to be. And some of those "model X is better than model Y" comparisons? Potentially measuring context pollution as much as model capability.&lt;/p&gt;

&lt;p&gt;I reckon we're going to see a wave of "actually, turns out our setup was fine, we just had too many tool schemas loaded" revelations over the next few months. The MCP ecosystem built a house of cards where adding more tools made every tool worse, and nobody measured it because the degradation was gradual.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Feels like:&lt;/strong&gt; Discovering your car's been running with the handbrake on for six months. The engine was fine the whole time. You were just dragging unnecessary weight.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; Anthropic built Tool Search because MCP's context cost is a real and measured problem. Fewer tools in context means better accuracy, full stop. If you're running more than a couple of MCP servers, either defer them or convert them. And if you want the full argument for why MCP is heading for the bin, the &lt;a href="https://dev.to/blog/death-of-mcp"&gt;Death of MCP&lt;/a&gt; piece lays it out end to end. The protocol had a good run. The ecosystem is moving on.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Peter Steinberger Says Just Talk To It, and He's Mostly Right</title>
      <dc:creator>Steven Gonsalvez</dc:creator>
      <pubDate>Fri, 10 Jul 2026 15:44:30 +0000</pubDate>
      <link>https://dev.to/stevengonsalvez/peter-steinberger-says-just-talk-to-it-and-hes-mostly-right-37bh</link>
      <guid>https://dev.to/stevengonsalvez/peter-steinberger-says-just-talk-to-it-and-hes-mostly-right-37bh</guid>
      <description>&lt;h2&gt;
  
  
  The Man Behind the Minimalism
&lt;/h2&gt;

&lt;p&gt;I met Peter Steinberger at Claude Anonymous in London sometime in June, and the presentation was properly fun. He's got that energy where you can tell he's not performing, he's just genuinely excited about the stuff he's building. Maintains a 300k line TypeScript React ecosystem. Web app, Chrome extension, CLI tool, Tauri desktop client, Expo mobile app. All maintained by one person with AI agents. And his setup is basically: open 3-8 parallel Codex instances, write short prompts (often 1-2 sentences plus a screenshot), let them go.&lt;/p&gt;

&lt;p&gt;He calls everything else "charade."&lt;/p&gt;

&lt;p&gt;I've been following his work since and I reckon he's one of the most practical voices in the agentic coding space. No hype. No frameworks for the sake of frameworks. Just a bloke who ships a lot of code with AI and has opinions about how to do it well.&lt;/p&gt;




&lt;h2&gt;
  
  
  His Principles (and Where I Stand)
&lt;/h2&gt;

&lt;p&gt;Here's the full list from his talk and blog posts. I agree with most of it. Where I don't, I'll say so.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Things I violently agree with:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Use tmux to run CLIs persistently.&lt;/strong&gt; Absolutely. This is foundational. If you're not doing this, start here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use &lt;a href="https://ast-grep.github.io/" rel="noopener noreferrer"&gt;ast-grep&lt;/a&gt; as a pre-commit hook.&lt;/strong&gt; Proper structural linting that catches things regex can't. Brilliant addition to any agent workflow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ask for options before executing.&lt;/strong&gt; I built a whole &lt;a href="https://github.com/stevengonsalvez/ainb-toolkit/tree/main/skills/interview" rel="noopener noreferrer"&gt;/interview skill&lt;/a&gt; around this. Get the model to present choices rather than guessing. Way better outcomes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Don't reset context. Cold start wastes time + tokens.&lt;/strong&gt; Agreed, but I'd go further. Don't reset, &lt;a href="https://github.com/stevengonsalvez/ainb-toolkit/tree/main/skills/handover" rel="noopener noreferrer"&gt;/handover&lt;/a&gt; instead. Summarise what was accomplished, carry it to the next session. Cold start is a waste. Context reset without transfer is nearly as bad.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Give examples.&lt;/strong&gt; Multi-shot prompting is always better than zero-shot. If you want the model to output something specific, show it what that looks like first. Massively improves one-shot accuracy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use voice input.&lt;/strong&gt; Absolutely. I started with &lt;a href="https://superwhisper.com/" rel="noopener noreferrer"&gt;Superwhisper&lt;/a&gt; and it changed how I work. Speaking your intent while pacing around the room is faster and more natural than typing. &lt;em&gt;(Update, February 2026: I've since switched to &lt;a href="https://dev.to/tools-tips/voice-coding"&gt;justspeaktoit&lt;/a&gt; by Chris Mitchelmore. Deepgram's $200 free credit and the speed is brilliant. Saved myself 9 quid a month on Wispr Flow.)&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Share screenshots.&lt;/strong&gt; Visual context for UI work is worth a thousand words of description. The model actually understands layouts from screenshots better than from your description of the layout.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write tests in the same context.&lt;/strong&gt; Tests written in the same session as the feature know the implementation intimately. Better coverage, catches bugs immediately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Schedule refactor time (20%).&lt;/strong&gt; Use &lt;a href="https://github.com/kucherenko/jscpd" rel="noopener noreferrer"&gt;jscpd&lt;/a&gt;, &lt;a href="https://knip.dev/" rel="noopener noreferrer"&gt;knip&lt;/a&gt;, &lt;a href="https://oxc.rs/" rel="noopener noreferrer"&gt;oxlint&lt;/a&gt;. Don't let debt accumulate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make small, atomic commit checkpoints.&lt;/strong&gt; Commit only what the agent touches. Easy to revert. Easy to review.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prefer CLI tools over MCP.&lt;/strong&gt; We wrote a &lt;a href="https://dev.to/blog/death-of-mcp"&gt;whole post about why MCP is a context tax&lt;/a&gt;. &lt;a href="https://dev.to/tools-tips/mcporter"&gt;mcporter&lt;/a&gt; for when you absolutely need MCP.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cancel long tasks and ask what's happening.&lt;/strong&gt; If the agent has been going for 5 minutes and you don't know what it's doing, stop it. Chances are it's in a hole.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep custom prompts minimal.&lt;/strong&gt; "commit" is a better prompt than "please create a well-formatted conventional commit message with a clear subject line." Less is more.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prefer Medium over High reasoning.&lt;/strong&gt; The model picks its own thinking depth. High reasoning burns tokens for marginal gains on most tasks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use 3-8 agents in parallel on a single repo.&lt;/strong&gt; Agree, though I'd cap at 4-5 max before the &lt;a href="https://dev.to/blog/workflow-methods-comparison"&gt;double pendulum problem&lt;/a&gt; kicks in.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Maintain an evolving AGENTS.md.&lt;/strong&gt; Delete stale guidelines. A lean AGENTS.md is worth more than a comprehensive one that's half wrong.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Where I disagree:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;"You don't need git worktrees."&lt;/strong&gt; Hard disagree. Worktrees are brilliant when agents need isolated branches. My swarm setup uses them heavily. The merge step adds overhead but the isolation prevents agents stomping on each other's files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Avoid subagents, RAG, etc."&lt;/strong&gt; Neutral on this. Subagents have real purposes: getting a code review done while you keep developing, running tests in parallel, preserving the main context window from getting polluted by research. The trick is knowing when the fresh context is worth the setup cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Prototype in a separate folder / PR."&lt;/strong&gt; Worktrees for this, not folders. Folders are messy. A worktree is a proper git branch with isolated state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"Avoid ALL-CAPS for Codex."&lt;/strong&gt; Fair for Codex. Claude Code is different. Sometimes emphasis matters in system prompts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The only real disagreement: instructions need structure.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Peter's philosophy is minimal instructions. Just talk to it. And for a solo dev on a single stack, that works. But the moment you're running multiple agents across different projects with varying conventions, you need guardrails. You need the agent to know "this project uses snake_case, that project uses camelCase" without you saying it every time. You need testing policies encoded, not remembered. You need the structure, even if you keep it lean.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Topology Insight
&lt;/h2&gt;

&lt;p&gt;The real variable isn't tools vs minimalism. It's project topology.&lt;/p&gt;

&lt;p&gt;Peter's workflow is a star: one developer, one codebase, all knowledge at the hub. That's 300k lines but it's singular. His approach is optimal for that shape.&lt;/p&gt;

&lt;p&gt;The moment you introduce parallel work streams (multiple repos, multiple agents needing different configs, team coordination), the topology becomes a mesh. Meshes need infrastructure that stars don't.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Geek Corner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Star vs mesh topology&lt;/strong&gt;: Steinberger's setup is one central developer with one codebase. Graph theorists call this a star. All decisions radiate from the hub. A team of developers running agents across multiple repos is a mesh: multiple actors, multiple codebases, shared context that needs synchronisation. The scaffold you need scales with the mesh density, not with some absolute measure of complexity. His advice is perfect for stars. It needs adaptation for meshes.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Feels like:&lt;/strong&gt; Two chefs arguing about whether you need a prep list. If you're cooking dinner for four, no. Crack on. If you're running a restaurant kitchen with six stations, yes. The answer depends on the kitchen, not the philosophy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; Peter's practices are solid regardless of your setup. Short prompts. Screenshots. Tests in context. Refactoring as 20% of your time. tmux. ast-grep. Voice input. These are good habits whether you wrap them in a framework or just talk to the model. I follow most of what he preaches. The disagreements are about scale and topology, not about the fundamentals. Read his stuff. Nick what works for you.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Ultrathink and Build: Weekly Dev Log with AI Tools and Side Projects</title>
      <dc:creator>Steven Gonsalvez</dc:creator>
      <pubDate>Fri, 10 Jul 2026 15:44:18 +0000</pubDate>
      <link>https://dev.to/stevengonsalvez/ultrathink-and-build-weekly-dev-log-with-ai-tools-and-side-projects-56le</link>
      <guid>https://dev.to/stevengonsalvez/ultrathink-and-build-weekly-dev-log-with-ai-tools-and-side-projects-56le</guid>
      <description>&lt;h2&gt;
  
  
  What I'm Reading and Watching This Week 📚
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;[Article]&lt;/strong&gt; "&lt;a href="https://example.com" rel="noopener noreferrer"&gt;Deep Work in the Age of AI&lt;/a&gt;" by Cal Newport. Proper interesting take on how AI coding tools change our relationship with focused work. Made me rethink a few things about how I structure my own deep work blocks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;[Video]&lt;/strong&gt; "The Art of Code Review" by ThePrimeagen. Some good bits on making code reviews actually useful instead of the usual rubber-stamping faff.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;[Book]&lt;/strong&gt; Currently reading "Staff Engineer" by Will Larson. If you're a senior dev wondering what comes next, this is the one.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Side Projects and AI Dev Tools I'm Building 🛠️
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;stevengonsalvez.com&lt;/strong&gt; 🟢 Active&lt;br&gt;
Next.js 15 + MDX blog platform with the byte-sized banter section you're reading right now. Foundation's done, blog section coming together. Next up is deploying to Vercel and sorting the domain.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Personal MCP Server&lt;/strong&gt; 🟡 On Hold&lt;br&gt;
Custom &lt;a href="https://dev.to/byte-sized-banter/week-28-mcp-context-kryptonite"&gt;MCP server&lt;/a&gt; for personal productivity tools. Paused while I wait for Claude Desktop to get its MCP support properly sorted.&lt;/p&gt;
&lt;h2&gt;
  
  
  What Changed This Week 🗞️
&lt;/h2&gt;

&lt;p&gt;Migrated the blog from dev.to-only to a self-hosted Next.js site. Launched this new weekly banter format. Still experimenting with how often I actually want to publish, reckon weekly is about right but we'll see.&lt;/p&gt;
&lt;h2&gt;
  
  
  Developer Tips and Tricks 💡
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick git alias for better logs:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git config &lt;span class="nt"&gt;--global&lt;/span&gt; alias.lg &lt;span class="s2"&gt;"log --graph --oneline --all --decorate"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This creates a beautiful visual git history that's way more readable than the default log.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude Desktop optimization tip:&lt;/strong&gt;&lt;br&gt;
Keep conversations small and restart often. The message limit resets every 5 hours, but shorter conversations use fewer tokens per message. You can get 2-3x more usage by starting fresh chats instead of continuing long threads. (If you're weighing up Claude Code versus other terminal tools, I did a &lt;a href="https://dev.to/byte-sized-banter/week-27-claude-vs-warp"&gt;proper comparison of Claude Code vs Warp AI&lt;/a&gt; a while back.)&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Claude Code Skills Just Made Half Your MCP Servers Redundant</title>
      <dc:creator>Steven Gonsalvez</dc:creator>
      <pubDate>Fri, 10 Jul 2026 15:44:05 +0000</pubDate>
      <link>https://dev.to/stevengonsalvez/claude-code-skills-just-made-half-your-mcp-servers-redundant-41d5</link>
      <guid>https://dev.to/stevengonsalvez/claude-code-skills-just-made-half-your-mcp-servers-redundant-41d5</guid>
      <description>&lt;h2&gt;
  
  
  Skills are what MCP should have been 📝
&lt;/h2&gt;

&lt;p&gt;Anthropic announced Agent Skills on October 16th, and I reckon this is one of those quiet releases that ends up mattering more than the flashy model drops.&lt;/p&gt;

&lt;p&gt;A skill is a markdown file with YAML frontmatter. That's it. No server process. No JSON-RPC. No WebSocket connections. No Docker containers. A markdown file. And the clever bit is how it loads.&lt;/p&gt;

&lt;p&gt;Level 1 is metadata. Name, description, trigger conditions. Costs you about 30 to 100 tokens and it's always in context. Level 2 is instructions. The actual "how to do the thing" content. Under 5,000 tokens. Only gets loaded when the model decides it's relevant to your task. Level 3 is resources. Referenced files, example code, whatever. Only pulled in when the instructions explicitly reference them.&lt;/p&gt;

&lt;p&gt;So you've got a system where 100 skills can sit in your context at a cost of maybe 3,000 to 10,000 tokens total. The metadata layer alone. The model reads the names and descriptions, figures out which ones matter, and loads only what it needs.&lt;/p&gt;

&lt;p&gt;Now compare that to MCP. Four or five MCP servers and you're looking at 40,000 to 60,000 tokens of JSON schemas. Loaded upfront. All of them. Whether you need them or not. Sitting in your context window like furniture in a flat you never use, taking up space and making the place harder to navigate.&lt;/p&gt;




&lt;h2&gt;
  
  
  Most MCP servers were never about live data
&lt;/h2&gt;

&lt;p&gt;Here's the thing I keep coming back to. What were people actually using MCP for?&lt;/p&gt;

&lt;p&gt;Some of it was legitimate live data connections. Database queries. API calls. Fetching real-time information the model can't have in its training data. Fair enough. That's a genuine use case and skills can't replace it.&lt;/p&gt;

&lt;p&gt;But a huge chunk of the MCP ecosystem was procedural knowledge. How to deploy this thing. How to format commits. How to run the test suite. How to interact with Jira. Step-by-step instructions wrapped in a protocol layer and loaded as tool definitions. Tens of thousands of tokens to tell the model "when someone asks about deployment, here's what to do."&lt;/p&gt;

&lt;p&gt;Skills do that at roughly 1/100th the context cost. Not an exaggeration. 30 tokens of metadata versus 3,000 tokens of tool schema. And the instructions only load when they're needed, so the actual runtime cost is even better than the comparison suggests.&lt;/p&gt;

&lt;p&gt;I've been running my own setup with skills for &lt;a href="https://dev.to/byte-sized-banter/week-33-terminal-wins"&gt;commits&lt;/a&gt;, code review, session management, research workflows, all sorts. Dozens of them. The context overhead is negligible. Try running dozens of MCP servers and watch your model forget what it was doing halfway through a conversation.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;📚 Geek Corner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Progressive loading vs. eager loading&lt;/strong&gt;: MCP uses eager loading. Every tool schema enters context at connection time. Skills use progressive loading across three tiers: metadata (always loaded, ~50 tokens each), instructions (loaded on relevance, &amp;lt;5k tokens), and resources (loaded on reference). The difference is architectural. MCP treats every tool as equally likely to be needed. Skills treat most tools as unlikely to be needed until proven otherwise. In information retrieval terms, MCP optimises for recall (everything available) while skills optimise for precision (only what's relevant). For context-constrained systems, precision wins.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Mario Zechner was asking the same question
&lt;/h2&gt;

&lt;p&gt;Two weeks after skills launched, Mario Zechner published "What if you don't need MCP at all?" on November 2nd. Different angle, same conclusion. His argument was that most of what people use MCP for can be handled by the agent calling CLIs and APIs directly, or by injecting knowledge through simpler mechanisms. Skills are exactly that simpler mechanism for the knowledge-injection half of the equation.&lt;/p&gt;

&lt;p&gt;I don't think this is subtle anymore. MCP tried to be the universal connector for everything: live data, procedural knowledge, tool access. Turns out that's too many jobs for one protocol, and the ecosystem is quietly unbundling it. Keep MCP for live data connections where you genuinely need them. Move procedural knowledge to skills. And for direct tool access, just call the CLI. Each replacement is cheaper and simpler than MCP was for that specific job.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Feels like:&lt;/strong&gt; Realising you've been driving a lorry to the corner shop. The lorry works, technically. But a bicycle gets you there faster, cheaper, and without having to find parking for a 12-tonne vehicle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; Skills are the right abstraction for procedural knowledge. MCP is the right abstraction for live data. Most people were using MCP for both and paying a massive context tax for the privilege. If you haven't moved your "how to do things" knowledge from MCP servers to skills, you're burning tokens for no reason. The &lt;a href="https://dev.to/blog/death-of-mcp"&gt;Death of MCP&lt;/a&gt; piece covers the full trajectory, but the short version is this: the protocol's scope is shrinking, and that's a good thing. Smaller scope, less overhead, better results. Crack on.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>AI Security Breaches, Vibe Coding Secrets Leak, and OpenAI's $500B Week</title>
      <dc:creator>Steven Gonsalvez</dc:creator>
      <pubDate>Fri, 10 Jul 2026 15:43:40 +0000</pubDate>
      <link>https://dev.to/stevengonsalvez/ai-security-breaches-vibe-coding-secrets-leak-and-openais-500b-week-51fo</link>
      <guid>https://dev.to/stevengonsalvez/ai-security-breaches-vibe-coding-secrets-leak-and-openais-500b-week-51fo</guid>
      <description>&lt;h2&gt;
  
  
  AI Security Disasters This Week 🔥
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Half a trillion in AI valuations. Eight thousand children's records nicked. Sudo is broken. Everything is on fire and we're shipping vibes.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Right. Where to start with this week of AI security breaches and vibe coding gone wrong.&lt;/p&gt;

&lt;p&gt;A ransomware gang called Radiant (love the branding, lads) hacked a nursery chain called Kido and walked off with personal data on 8,000 children. &lt;em&gt;Children.&lt;/em&gt; Not enterprise accounts. Not crypto wallets. Actual kids in actual nurseries. I don't usually get properly miffed about security news because frankly if you're still leaving RDP open to the internet you deserve what's coming. But targeting nurseries? That's a new kind of grim.&lt;/p&gt;

&lt;p&gt;Meanwhile: OpenAI hit a $500 billion valuation. Half. A. Trillion. We're living in a timeline where AI companies are worth more than most countries' GDP and a nursery chain can't keep toddler data safe. Cool. This is fine.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Hits Keep Coming 💀
&lt;/h2&gt;

&lt;p&gt;It wasn't just the nursery hack. This week was a proper security shambles from top to bottom.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Google Ads serving trojans.&lt;/strong&gt; You search for something legitimate, click an ad at the top of Google, and congratulations, you've just installed malware. Google taking money to distribute malware is &lt;em&gt;chef's kiss&lt;/em&gt; levels of ironic. The ad platform that prints money can't vet what it's printing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fake invoices spreading RATs.&lt;/strong&gt; Remote Access Trojans shipped via invoice PDFs. Because apparently we still haven't sorted out "don't open random attachments" after twenty-odd years of trying.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The sudo exploit&lt;/strong&gt; (CVE-2025-32463) got added to CISA's Known Exploited Vulnerabilities list. Actively exploited. In the wild. Right now. Sudo. The thing that literally gates root access on every Linux box you've ever touched. If that doesn't make you sweat a bit, you're not paying attention.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tile tracking devices&lt;/strong&gt; flagged as a stalking risk. The thing you bought to find your keys can apparently be weaponised to find &lt;em&gt;you&lt;/em&gt;. Reassuring.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;UK Co-Op attack costs hit $275 million.&lt;/strong&gt; DragonForce's April attack resulted in weeks of empty shelves and a quarter billion in damages. That's not a cyber incident. That's an economic event.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Silent smishing&lt;/strong&gt;, SMS phishing that doesn't even trigger a notification. Your phone gets compromised and you don't even know it happened. Brilliant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Feels like:&lt;/strong&gt; Living in a horror film where every door you open has something worse behind it, and someone in the background keeps cheerfully announcing record-breaking funding rounds.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Vibe Coding Security Reckoning 🫠
&lt;/h2&gt;

&lt;p&gt;Here's where it gets properly interesting. "Vibe coded secrets leak" was an actual headline this week. Apps built by AI, people just accepting whatever the model spits out, shipping with hardcoded API keys and credentials sitting right there in the code.&lt;/p&gt;

&lt;p&gt;I've said it before and I'll keep banging this drum: vibe coding is sick for prototyping. It's magic for scaffolding. But the moment you ship vibe-coded output without reviewing it, you're essentially deploying code that &lt;em&gt;nobody&lt;/em&gt; has read. Not you. Not the model (it doesn't &lt;em&gt;read&lt;/em&gt; code, it predicts tokens). Nobody.&lt;/p&gt;

&lt;p&gt;And now those predicted tokens include your AWS keys. Outstanding.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;📚 Geek Corner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;The secrets-in-AI-code problem&lt;/strong&gt;: When an LLM generates code, it draws on patterns from training data, which includes thousands of tutorials and Stack Overflow answers with placeholder credentials. The model doesn't &lt;em&gt;know&lt;/em&gt; those are supposed to be placeholders. It's pattern-matching, and the pattern is "config file has a string here." Combine that with developers who treat AI output as trusted input, skip code review, and push straight to main... you get secrets in production. The fix isn't to stop using AI for code. It's to treat AI output the same way you'd treat a PR from an intern: review everything, run secret scanning (git-secrets, trufflehog, gitleaks), and never trust generated config values. The tooling exists. People just aren't using it because vibes.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Meanwhile, In AI Utopia 🚀
&lt;/h2&gt;

&lt;p&gt;While the security world was having a week, the AI hype machine was running at full chat.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Anthropic dropped Claude Sonnet 4.5&lt;/strong&gt; along with the &lt;strong&gt;Claude Agent SDK&lt;/strong&gt;. The model topped SWE-bench Verified, which means it's now the best at fixing code that probably got compromised because someone else's AI wrote dodgy code. The circle of life. (I put Claude Code through its paces &lt;a href="https://dev.to/byte-sized-banter/week-41-ios-simulator-showdown"&gt;against Codex on an iOS simulator build&lt;/a&gt; the following week, if you're curious how it actually performs.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OpenAI launched Sora 2&lt;/strong&gt;, photorealistic video generation where you can insert yourself into generated footage. Nothing concerning about deepfake technology going mainstream the same week we're discussing social engineering attacks. Nothing at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OpenAI hit $500 billion valuation&lt;/strong&gt; after a $6.6 billion secondary share sale. The world's most valuable startup. I don't even know what to do with that number. The Co-Op hack cost $275 million. OpenAI's valuation could absorb roughly 1,800 Co-Op-scale attacks and still be worth something. The scales are broken.&lt;/p&gt;

&lt;p&gt;And in the developer corner: &lt;strong&gt;comment-driven development&lt;/strong&gt; became a talking point, the idea that since LLMs rely on comments, well-commented code is now functional documentation for your AI pair programmer. Also, someone did a proper write-up on &lt;strong&gt;Claude Code's magic&lt;/strong&gt; and how its agentic patterns actually work under the hood. Both good reads, both slightly surreal to be discussing while sudo is actively exploited.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Hiring Hot Take 🌶️
&lt;/h2&gt;

&lt;p&gt;The Changelog ran a piece this week: &lt;strong&gt;"Hiring only senior engineers is killing companies."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I've got thoughts on this one. The argument goes that if you only hire seniors, you end up with a team of people who are all good at making decisions but nobody wants to do the actual building. Everyone's an architect, nobody's laying bricks.&lt;/p&gt;

&lt;p&gt;There's truth in it. I've seen teams of eight seniors spend three sprints debating an API schema that a motivated mid-level would've shipped in a week. But the counter-argument, that you need juniors to do the grunt work, is properly outdated now. The grunt work is increasingly what AI handles. The boring CRUD endpoints, the boilerplate, the scaffolding. That was the junior dev pipeline, and it's being automated.&lt;/p&gt;

&lt;p&gt;So where does that leave us? Reckon the real problem isn't "too many seniors" but "too many people who only know how to be senior in the old way." The game's changing. Being senior used to mean you'd seen every pattern. Now it means you can evaluate whether the pattern the AI suggested is actually right, or whether it's about to ship your secrets to GitHub.&lt;/p&gt;

&lt;p&gt;Full circle, innit.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Bit That Keeps Me Up 😶
&lt;/h2&gt;

&lt;p&gt;Here's what's actually bothering me about this week. It's not any single story. It's the &lt;em&gt;gap&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;On one side: AI companies hitting half-trillion valuations. Agent SDKs launching. Photorealistic video generation. Comment-driven development. The future is here and it's magnificent.&lt;/p&gt;

&lt;p&gt;On the other side: nurseries getting hacked. Sudo exploits in the wild. Vibe-coded apps leaking credentials. Invoice phishing still working. The fundamentals are on fire and nobody's watching because everyone's distracted by the shiny new model release.&lt;/p&gt;

&lt;p&gt;We're building the most sophisticated software in human history on top of infrastructure that can't even keep nursery data safe. And the response from the industry is to ship faster, review less, and let the AI handle it.&lt;/p&gt;

&lt;p&gt;I'm not saying slow down. I'm saying &lt;em&gt;look down&lt;/em&gt;. The foundations are cracking and we're too busy adding floors.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; Week 40 was a horror show wearing a party hat. Record valuations upstairs, ransomware in the basement. The AI gold rush is real, but so are the 8,000 kids whose data is floating around a dark web forum right now. Maybe, just maybe, we should spend as much energy on the boring security stuff as we do on the next model benchmark. But we won't. Because vibes.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Multi-Agent Error Cascades: The Double Pendulum Problem Nobody Talks About</title>
      <dc:creator>Steven Gonsalvez</dc:creator>
      <pubDate>Fri, 10 Jul 2026 15:43:27 +0000</pubDate>
      <link>https://dev.to/stevengonsalvez/multi-agent-error-cascades-the-double-pendulum-problem-nobody-talks-about-1o22</link>
      <guid>https://dev.to/stevengonsalvez/multi-agent-error-cascades-the-double-pendulum-problem-nobody-talks-about-1o22</guid>
      <description>&lt;h2&gt;
  
  
  Your agents are a chaos machine 🎯
&lt;/h2&gt;

&lt;p&gt;So HumanLayer dropped their Advanced Context Engineering piece last week, and buried in it is this absolute banger of a line:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Bad research -&amp;gt; bad plan -&amp;gt; bad code. A single wrong line in research cascades to widespread errors."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's Dex Horthy describing what happens when you chain AI agents together without checkpoints. And it's the exact same mechanics as a double pendulum.&lt;/p&gt;

&lt;p&gt;You know the double pendulum, yeah? Simple physics demo. One pendulum hanging from another. The top one swings predictably. The bottom one goes absolutely mental. Tiny changes in the initial swing of the top pendulum produce wildly different trajectories in the bottom one. Chaos theory in action. Looks like it's possessed.&lt;/p&gt;

&lt;p&gt;Multi-agent AI systems are double pendulums.&lt;/p&gt;

&lt;p&gt;Your research agent makes a small mistake. Misidentifies which module handles authentication. Gets one import path wrong. Reads an outdated API signature. Tiny error. Barely noticeable in the research output.&lt;/p&gt;

&lt;p&gt;Your planning agent reads that research and builds a plan around the wrong assumption. The plan isn't obviously wrong. It's coherent. It just starts from a slightly incorrect foundation. The deviation from reality is bigger now but still plausible-looking.&lt;/p&gt;

&lt;p&gt;Your implementation agent reads that plan and writes hundreds of lines of code. Code that's internally consistent but built on a foundation of sand. The cascade is complete. One wrong line in research became a hundred wrong lines in code.&lt;/p&gt;




&lt;h3&gt;
  
  
  Why more agents makes this worse, not better
&lt;/h3&gt;

&lt;p&gt;Here's the bit that does my head in. The whole pitch of multi-agent systems is "more agents = better results." Specialise each agent. Divide the labour. Sounds reasonable.&lt;/p&gt;

&lt;p&gt;But every agent you add to the chain is another joint in the pendulum. Every handoff is another point where small errors amplify. A three-agent pipeline (research, plan, implement) has two handoff points. A five-agent pipeline has four. Each one is a potential chaos amplification.&lt;/p&gt;

&lt;p&gt;Sean Moran documented this properly in January 2026, showing that unstructured multi-agent architectures amplify errors 17x compared to single-agent baselines. 17x! You're not getting better results by adding agents. You're getting 17x worse results if the coordination is sloppy.&lt;/p&gt;

&lt;p&gt;This is why HumanLayer's RPI methodology forces human review between phases. It's not because humans are better at writing code (the agents have us beat there). It's because humans are circuit breakers. They catch the small research error before it cascades into hundreds of bad code lines. The human doesn't need to be a better researcher than the agent. They just need to spot when the research is off.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;📚 Geek Corner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Why the double pendulum analogy is more than metaphor.&lt;/strong&gt; In dynamical systems theory, sensitivity to initial conditions means that prediction accuracy degrades exponentially with each degree of freedom added. A single pendulum is predictable forever. A double pendulum is predictable for a short time, then diverges. A triple pendulum is basically random. Multi-agent AI chains follow the same mathematical pattern. Each agent's output has some error distribution. When that output feeds into the next agent, the errors don't add linearly. They multiply. Agent 1's 5% error rate doesn't become 10% at agent 2. It becomes 5% of correct paths times agent 2's own error rate on each, which is a branching tree of possible failures. The paper "Agents of Chaos" (February 2026, arxiv) formalises this: competing prediction agents can phase-transition from stability into mathematical chaos. This isn't a soft analogy. It's the same maths.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h3&gt;
  
  
  The fix is boring (and that's the point)
&lt;/h3&gt;

&lt;p&gt;Human checkpoints between agent phases. That's it. That's the fix.&lt;/p&gt;

&lt;p&gt;HumanLayer calls it RPI (Research, Plan, Implement) and the whole point is that a human reviews 400 lines of research and plan before the agent writes 4,000 lines of code. You're spending 10 minutes of human review to prevent 10 hours of debugging broken output.&lt;/p&gt;

&lt;p&gt;Not every agent chain needs human checkpoints. A two-agent system doing research and summarisation? The pendulum doesn't have enough joints to go chaotic. But the moment you've got three or more agents in a chain, with each one's output feeding the next one's input, you're playing double pendulum roulette. Add a human circuit breaker at the highest-leverage handoff point (usually between research and implementation) and the chaos stays contained.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Feels like:&lt;/strong&gt; Stacking Jenga blocks on top of a washing machine. Each individual block is stable. The tower is not. The fix isn't better blocks. It's checking the tower every few layers.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; More agents is not better agents. Every handoff is a potential chaos amplification point. If you're running multi-agent pipelines without human checkpoints between phases, you're building a double pendulum and hoping it doesn't go chaotic. Spoiler: it will. HumanLayer's RPI is one answer. Any answer that puts a human circuit breaker at the right handoff point works. The expensive mistake is skipping the checkpoint, not adding it.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Why Terminal AI Coding Agents Are Beating IDE Extensions</title>
      <dc:creator>Steven Gonsalvez</dc:creator>
      <pubDate>Fri, 10 Jul 2026 15:43:14 +0000</pubDate>
      <link>https://dev.to/stevengonsalvez/why-terminal-ai-coding-agents-are-beating-ide-extensions-1ipo</link>
      <guid>https://dev.to/stevengonsalvez/why-terminal-ai-coding-agents-are-beating-ide-extensions-1ipo</guid>
      <description>&lt;h2&gt;
  
  
  The Headless Takeover 🖥️
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;IDEs had a good run. Then the terminals learned to think.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Here's the thing nobody's saying out loud: terminal-based coding agents are eating headed apps for breakfast. And it's not even close.&lt;/p&gt;




&lt;h3&gt;
  
  
  The Architecture Gap
&lt;/h3&gt;

&lt;p&gt;Every IDE-based agent - Cursor, Windsurf, whatever's launching next week - has the same fundamental constraint: &lt;strong&gt;IPC&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Inter-process communication. The IDE runs here, the AI runs there, and they talk through some protocol layer. Extensions, language servers, message passing. It works, but it's a bottleneck by design.&lt;/p&gt;

&lt;p&gt;Terminal agents? &lt;strong&gt;Direct subprocess execution.&lt;/strong&gt; No middleware. No protocol translation. You spawn a process, it runs, you get output. Done.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;📚 Geek Corner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;IPC vs Subprocess&lt;/strong&gt;: IDE extensions typically communicate via JSON-RPC, WebSockets, or custom protocols. Each message serialises, transmits, deserialises. Terminal agents skip all of that - they're just running shell commands as child processes with direct stdin/stdout pipes. The overhead difference is orders of magnitude.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h3&gt;
  
  
  The Composability Factor
&lt;/h3&gt;

&lt;p&gt;Terminal tools are &lt;em&gt;composable by default&lt;/em&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cat &lt;/span&gt;file.ts | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="s2"&gt;"TODO"&lt;/span&gt; | &lt;span class="nb"&gt;wc&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Thirty years of Unix philosophy baked into every interaction. Pipes. Redirects. Scripts. You can wrap &lt;em&gt;anything&lt;/em&gt; - linters, formatters, compilers, test runners, deployment scripts - and the agent just treats them as tools.&lt;/p&gt;

&lt;p&gt;Try doing that in a GUI. You're clicking buttons. Waiting for panels to load. Fighting with extension APIs that change every release.&lt;/p&gt;

&lt;p&gt;Terminal agents don't care about your UI framework. They run commands. Commands compose. That's it.&lt;/p&gt;




&lt;h3&gt;
  
  
  The Scaling Play
&lt;/h3&gt;

&lt;p&gt;This is where it gets spicy.&lt;/p&gt;

&lt;p&gt;How do you scale a GUI-based agent? You don't. One human, one screen, one IDE instance. Maybe you run a second window if you're feeling fancy.&lt;/p&gt;

&lt;p&gt;Terminal agents? Spawn ten of them. Spawn a hundred. Run them in CI. Run them on every PR. Run them overnight while you sleep. They're just processes - orchestrate them however you want.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# This is trivial with terminal agents&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;repo &lt;span class="k"&gt;in &lt;/span&gt;repos/&lt;span class="k"&gt;*&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;claude &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"Run tests and fix failures"&lt;/span&gt; &lt;span class="nt"&gt;--dir&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$repo&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &amp;amp;
&lt;span class="k"&gt;done
&lt;/span&gt;&lt;span class="nb"&gt;wait&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Try automating that with Cursor.&lt;/p&gt;




&lt;h3&gt;
  
  
  The Browser Delusion
&lt;/h3&gt;

&lt;p&gt;"But browser-based agents can see the screen! They can click buttons!"&lt;/p&gt;

&lt;p&gt;Sure. And they're slow, brittle, and break every time someone changes a CSS class. Screen scraping is not a foundation - it's a hack.&lt;/p&gt;

&lt;p&gt;Terminal agents don't need to &lt;em&gt;see&lt;/em&gt; your app. They read logs. They query APIs. They run commands that return structured data. Deterministic, scriptable, testable.&lt;/p&gt;

&lt;p&gt;The browser-based crowd is building on sand. The terminal crowd is building on decades of battle-tested infrastructure.&lt;/p&gt;




&lt;h3&gt;
  
  
  What's Coming
&lt;/h3&gt;

&lt;p&gt;This is just the beginning.&lt;/p&gt;

&lt;p&gt;Right now we're running single agents on single tasks. But the terminal architecture enables something bigger: &lt;strong&gt;agent swarms&lt;/strong&gt;. Parallel execution across codebases. Agents spawning agents. Hierarchical task decomposition with actual process isolation.&lt;/p&gt;

&lt;p&gt;None of that works if your agent needs a display. All of it works when your agent is just a process.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; Headed agents are demos. Terminal agents are infrastructure. The market hasn't figured this out yet, but it will. If you want to see this play out in practice, the &lt;a href="https://dev.to/byte-sized-banter/week-27-claude-vs-warp"&gt;Claude Code vs Warp showdown&lt;/a&gt; is a good place to start. And the &lt;a href="https://dev.to/byte-sized-banter/week-29-vibe-coding-peak-hype"&gt;vibe coding peak hype&lt;/a&gt; week showed exactly why the IDE approach is running into walls.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>GPT-5, Opus 4.1, and Duct-Tape Security: AI's Wildest Week in 2025</title>
      <dc:creator>Steven Gonsalvez</dc:creator>
      <pubDate>Fri, 10 Jul 2026 15:43:02 +0000</pubDate>
      <link>https://dev.to/stevengonsalvez/gpt-5-opus-41-and-duct-tape-security-ais-wildest-week-in-2025-3fhi</link>
      <guid>https://dev.to/stevengonsalvez/gpt-5-opus-41-and-duct-tape-security-ais-wildest-week-in-2025-3fhi</guid>
      <description>&lt;h2&gt;
  
  
  The Week That Had Everything 🎪
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;A 24-year-old lands a $250M AI pay package. Meanwhile, link wrappers are nicking your login credentials. Same industry. Same week.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Right, where do I even start with this one. Week 32 was the kind of week where every newsletter hit different. GPT-5 dropped. Opus 4.1 dropped. OpenAI went open-source. Cursor got a CLI &lt;em&gt;and&lt;/em&gt; got poisoned. North Korean devs are still out here catfishing hiring managers. And someone, somewhere, is writing a quarter-billion-dollar cheque to a researcher who can't legally rent a car in most US states.&lt;/p&gt;

&lt;p&gt;Let's crack on.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Money's Gone Absolutely Mental 💰
&lt;/h2&gt;

&lt;p&gt;AI researchers are being recruited like Premier League strikers now. We're talking $250M packages. For a 24-year-old. I don't care how good your transformer architecture paper is, that number should make everyone uncomfortable.&lt;/p&gt;

&lt;p&gt;The maths works out to roughly "we'd rather overpay by 10x than let a competitor have you." Which, fine, that's how bidding wars work. But it tells you something about the state of things when the talent pool is so thin that a single researcher commands more than most companies are &lt;em&gt;worth&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Here's what bothers me though. These packages aren't salary. They're structured as equity, retention bonuses, and golden handcuffs. The researcher doesn't actually get $250M unless the company's valuation holds. And if we've learned anything from the last two decades of tech, valuations are vibes until they're not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Feels like:&lt;/strong&gt; Paying someone a quarter billion to fix your plumbing while the rest of the house is on fire and nobody's rung the fire brigade.&lt;/p&gt;




&lt;h2&gt;
  
  
  GPT-5: The Main Event (That Got Upstaged) 🎬
&lt;/h2&gt;

&lt;p&gt;OpenAI shipped GPT-5 on Friday. Three variants: Pro, Mini, and Nano. Available to everyone in ChatGPT. Smartest and fastest, they say.&lt;/p&gt;

&lt;p&gt;And honestly? The reaction was a bit muted. Not because GPT-5 is bad. By all accounts it's proper good. But it landed in a week where &lt;em&gt;everything&lt;/em&gt; dropped. Opus 4.1 on Wednesday. GPT-OSS on Wednesday. Gemini coding agent. Cursor CLI. The news cycle was so saturated that a flagship model launch felt like just another item on the list.&lt;/p&gt;

&lt;p&gt;That's either a sign of how fast things move now, or a sign that we've collectively lost the ability to be impressed for more than about four hours.&lt;/p&gt;

&lt;p&gt;The GPT-OSS release is actually the more interesting story, if you ask me. A 120B open-weight model under Apache 2.0 that nearly matches o4-mini on reasoning benchmarks and runs on a single 80GB GPU. OpenAI going &lt;em&gt;actually&lt;/em&gt; open-source after years of "open" being a punchline. DeepSeek and Qwen3 still beat it on raw intelligence, but the fact that OpenAI is playing this game at all is a shift.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;📚 Geek Corner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;GPT-OSS-120B and the open-weight arms race&lt;/strong&gt;: OpenAI's 120B model hits near-parity with o4-mini on core reasoning while fitting on a single A100/H100. The trick is aggressive distillation from their larger proprietary models, not architectural novelty. This puts it behind DeepSeek R1 and Qwen3-235B on intelligence benchmarks, but the Apache 2.0 licence means it's actually deployable without lawyers. The real question: does OpenAI releasing competitive open models undermine their own API revenue, or does it function as a loss leader to keep developers in the OpenAI ecosystem? My bet is the latter. Get them building on your architecture, then upsell the proprietary stuff. Classic developer relations playbook, just at a different scale.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Opus 4.1: The Quiet Mid-Week Drop 🔬
&lt;/h2&gt;

&lt;p&gt;Anthropic shipped Claude Opus 4.1 on Wednesday. Improvements in agentic tasks, real-world coding, and reasoning. No massive fanfare. No countdown timer. Just... here it is, it's better, carry on.&lt;/p&gt;

&lt;p&gt;I've been running Claude Code daily and the improvements in agentic task completion are noticeable. Less faff with multi-step operations. Better at holding context across long sessions. The kind of upgrade that doesn't make you say "wow" but does make you say "huh, that worked first time" more often.&lt;/p&gt;

&lt;p&gt;Multiple newsletters called it "the biggest AI week of the year" and they weren't wrong. Three major model releases in a single week from three different companies. Proper arms race energy.&lt;/p&gt;




&lt;h2&gt;
  
  
  Meanwhile, Everything Is On Fire 🔥
&lt;/h2&gt;

&lt;p&gt;Right, so while everyone's mucking about with their shiny new models, let's talk about the absolute state of security this week.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Link wrappers stealing logins.&lt;/strong&gt; Cloudflare's Email Security team caught threat actors abusing link-wrapping services from Proofpoint and Intermedia. The tools that are meant to &lt;em&gt;protect&lt;/em&gt; you from dodgy links are being weaponised to &lt;em&gt;deliver&lt;/em&gt; dodgy links. That's not a vulnerability, that's a comedy sketch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Public prompts to local shells.&lt;/strong&gt; Exactly what it sounds like. Prompts that can escape into your local shell. If you're running AI tools that execute code, this should keep you up at night.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cursor MCPoison.&lt;/strong&gt; The same week Cursor launches their CLI, a poisoning attack surfaces that targets Cursor specifically. You couldn't write better timing if you tried. Build the tool on Monday, someone finds a way to poison it by Friday.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;North Korean fake devs.&lt;/strong&gt; Still happening. Thousands of IT workers deployed abroad with fake identities, landing remote jobs at Western companies. We've known about this for over a year and the industry response has been... well, it's been nothing much, hasn't it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Perplexity's stealth crawlers.&lt;/strong&gt; Cloudflare caught Perplexity using stealth crawling to bypass website restrictions. Robots.txt? Never heard of her. We're just going to hoover up your content and serve it back without attribution. Cheeky doesn't begin to cover it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Chrome cookie encryption blown up. Student financial data stolen. Informants exposed in a hack. Summer cyber attack spike.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;All the same week. All happening while someone signs a $250M retention package.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Smell of Vibe Coding 😷
&lt;/h2&gt;

&lt;p&gt;Changelog ran a piece titled "The smell of vibe coding" and I proper love the framing. Because vibe coding &lt;em&gt;does&lt;/em&gt; have a smell. It's the smell of code that works but nobody understands why. It's the smell of a codebase where the AI wrote 80% of it and the human vibed their way through the remaining 20%. It compiles. Tests pass. Ship it.&lt;/p&gt;

&lt;p&gt;Until it doesn't. And nobody can debug it because nobody actually &lt;em&gt;wrote&lt;/em&gt; it.&lt;/p&gt;

&lt;p&gt;The Gemini coding agent also launched this week, because apparently every company needed to ship something. Google's approach is different to Claude Code and Cursor, leaning harder into the "agent that does things for you" model rather than the "copilot that helps you do things" model. Jury's still out on which philosophy wins.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;📚 Geek Corner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;The coding agent spectrum&lt;/strong&gt;: There's a real philosophical split forming. On one end: Claude Code and Cursor, which are essentially power tools. You drive, they assist. On the other: Gemini's coding agent and similar products that try to do the whole job autonomously. The first approach keeps the developer in the loop but limits throughput. The second approach scales but introduces the "smell" problem. Nobody knows how to maintain code they didn't write, whether the author was a junior dev in Bangalore or an LLM in a data centre. The answer is probably somewhere in the middle, but right now everyone's racing to the extremes.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  The Juxtaposition That Won't Leave My Head 🤯
&lt;/h2&gt;

&lt;p&gt;Here's what I keep coming back to.&lt;/p&gt;

&lt;p&gt;On Monday, someone signs a $250M deal because AI talent is &lt;em&gt;that&lt;/em&gt; valuable.&lt;/p&gt;

&lt;p&gt;On Tuesday, North Korean operatives are infiltrating Western companies with fake dev profiles.&lt;/p&gt;

&lt;p&gt;On Wednesday, three flagship AI models drop simultaneously.&lt;/p&gt;

&lt;p&gt;On Thursday, Chrome's cookie encryption gets cracked and student financial data gets nicked.&lt;/p&gt;

&lt;p&gt;On Friday, GPT-5 launches alongside a poisoning attack on one of the most popular AI coding tools.&lt;/p&gt;

&lt;p&gt;We're building absurdly capable systems and securing them with bodge jobs and hope. The same industry that can afford quarter-billion retention packages can't figure out how to stop phishing attacks that abuse its own security tools. The gap between what we're &lt;em&gt;building&lt;/em&gt; and how well we're &lt;em&gt;protecting&lt;/em&gt; it is getting wider every week.&lt;/p&gt;

&lt;p&gt;And Perplexity's over there just crawling whatever it wants, because apparently rules are for other people.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's Actually Worth Your Time This Week
&lt;/h2&gt;

&lt;p&gt;If you only track three things from this week:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;GPT-OSS under Apache 2.0&lt;/strong&gt; - More interesting than GPT-5 itself. Open-weight models that rival proprietary ones change the economics for everyone running inference at scale.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Cursor MCPoison&lt;/strong&gt; - If you're using AI coding tools (and you are), this is the attack vector you need to understand. Poisoned context that makes your AI write vulnerable code. Sneaky and nasty.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The link-wrapper abuse&lt;/strong&gt; - Security tools being turned into attack vectors. If your org uses Proofpoint for email protection, go check your configs.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; We're in the era where a single week delivers three frontier model launches, an open-source bombshell, and half a dozen security disasters. The money's flowing. The models are shipping. The security is held together with duct tape and wishful thinking. I wrote about &lt;a href="https://dev.to/blog/mcp-ai-security-risks-best-practices"&gt;MCP security risks in detail&lt;/a&gt; earlier this year, and the Cursor MCPoison story is exactly the kind of thing I was warning about. If you reckon &lt;a href="https://dev.to/byte-sized-banter/week-29-vibe-coding-peak-hype"&gt;vibe coding peaked last month&lt;/a&gt;, this week says hold my beer. If that doesn't sum up 2025, I don't know what does.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Vibe Coding Peak Hype: Windsurf Acquisition Chaos and the AI IDE Wars</title>
      <dc:creator>Steven Gonsalvez</dc:creator>
      <pubDate>Fri, 10 Jul 2026 15:42:49 +0000</pubDate>
      <link>https://dev.to/stevengonsalvez/vibe-coding-peak-hype-windsurf-acquisition-chaos-and-the-ai-ide-wars-pfi</link>
      <guid>https://dev.to/stevengonsalvez/vibe-coding-peak-hype-windsurf-acquisition-chaos-and-the-ai-ide-wars-pfi</guid>
      <description>&lt;h2&gt;
  
  
  Who Actually Bought Windsurf? 🌀
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Three companies walk into a bar. They all claim they bought the same startup.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Right. So here's the week in AI coding acquisitions, and I need you to stay with me because it gets properly daft.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Monday&lt;/strong&gt;: OpenAI's $3 billion bid for Windsurf collapses after Anthropic yanks their Claude API access. Brutal move, that. Like cutting off your rival's electricity mid-negotiation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Also Monday&lt;/strong&gt;: Google DeepMind swoops in and hires Windsurf's CEO Varun Mohan, co-founder Douglas Chen, and the key researchers. Not an acquisition. A reverse acqui-hire. Price tag: $2.4 billion. For the &lt;em&gt;people&lt;/em&gt;, mind you. Not the product.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tuesday&lt;/strong&gt;: Cognition (the Devin lot) scoops up what's left: 250 engineers and $82 million in ARR.&lt;/p&gt;

&lt;p&gt;So in the space of 48 hours, Windsurf got rejected, stripped for parts, and sold off at a car boot sale. If you're a Windsurf user wondering what's happening to your IDE, the honest answer is: nobody has a clue, least of all Windsurf.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Feels like:&lt;/strong&gt; Watching a pub quiz team implode mid-game, with three other teams fighting over which one gets to adopt the remaining members.&lt;/p&gt;




&lt;h2&gt;
  
  
  Steve Yegge Says What Everyone's Thinking
&lt;/h2&gt;

&lt;p&gt;The Pragmatic Engineer ran an interview with Steve Yegge this week, you know, the bloke who's been ranting about platforms since before most current SWEs had GitHub accounts.&lt;/p&gt;

&lt;p&gt;His take on vibe coding: &lt;em&gt;it's deceptively hard&lt;/em&gt;. Not the coding itself (the AI handles that bit). The hard part is knowing whether what the AI produced is any good. He reckons there's an emerging "AI Fixer" role inside companies, someone whose entire job is reviewing and correcting AI-generated code.&lt;/p&gt;

&lt;p&gt;Which... yeah. That tracks. I've been saying this for months. Vibe coding is mint when you're prototyping or bodging together a script. But the second you need it to work in production, at scale, with actual users? You need someone who &lt;em&gt;understands&lt;/em&gt; the code the machine spat out. And right now that someone is still a human.&lt;/p&gt;

&lt;p&gt;The hot take from the same week, "all models are the same", is doing the rounds on tech Twitter. And I reckon it's about 60% correct. For most tasks, swapping Claude for GPT for Gemini gets you roughly comparable results. The differentiation is in the tooling, the UX, the ecosystem. Not the raw model output.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;📚 Geek Corner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;The "AI Fixer" pattern&lt;/strong&gt;: Yegge's observation maps onto something systems thinkers have known for ages. Every wave of automation creates a new class of worker whose job is managing the automation's mistakes. Assembly lines created quality inspectors. Automated trading created risk managers. AI coding is creating AI fixers. The interesting bit is where these fixers sit: inside the IDE (like Cursor's diff review), in CI (automated code review bots), or as an actual human role. My bet: all three, layered. The tooling will catch the obvious stuff, CI will catch the structural stuff, and humans will catch the "this technically works but is completely wrong for our domain" stuff.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  AWS Kiro: Another IDE Enters the Chat
&lt;/h2&gt;

&lt;p&gt;Amazon launched Kiro this week. An "agentic IDE" because apparently we needed another one of those.&lt;/p&gt;

&lt;p&gt;Look, I haven't given it a proper go yet so I'll reserve judgment. But the timing is... interesting. Windsurf just got dismembered. Cursor is the darling of the vibe coding crowd. VS Code with Copilot is the safe enterprise pick. And now AWS thinks what the market needs is &lt;em&gt;their&lt;/em&gt; IDE?&lt;/p&gt;

&lt;p&gt;The pitch is all about "spec-driven development", where you write a spec, Kiro builds it. Which sounds ace until you remember that writing a good spec is &lt;em&gt;the hard part of software engineering&lt;/em&gt;. We didn't solve that problem when we had humans writing code, and bolting an AI on top doesn't make the requirements magically clearer.&lt;/p&gt;

&lt;p&gt;Still. Amazon has a habit of shipping things that look naff at launch and then quietly becoming essential three years later. See: Lambda, ECS, CDK. So maybe this time next year I'll be eating my words.&lt;/p&gt;




&lt;h2&gt;
  
  
  China Drops a Trillion-Parameter Model (For Free)
&lt;/h2&gt;

&lt;p&gt;While everyone in the West is arguing about who owns Windsurf, China quietly dropped a trillion-parameter open model. Free to use. Free to fine-tune.&lt;/p&gt;

&lt;p&gt;The "AI is a commodity" crowd is having a field day. And honestly? The commodity argument gets stronger every week. When China can ship models this big, open-source, at no cost, what exactly is the moat for commercial model providers?&lt;/p&gt;

&lt;p&gt;I'll tell you what it is: integration and trust. Nobody's running a trillion-parameter Chinese model on their healthcare data or financial systems. Not because the model isn't capable, but because compliance, data residency, and the fact that your CISO will have an actual cardiac event.&lt;/p&gt;

&lt;p&gt;But for hobbyists, researchers, and anyone mucking about with prototypes? Absolute gift. The walls around premium AI are getting shorter by the month.&lt;/p&gt;




&lt;h2&gt;
  
  
  Claude Code: The Main Character of Week 29
&lt;/h2&gt;

&lt;p&gt;Every other newsletter this week had a Claude Code tips article. Claude Code Router (route your requests to different models). Claude Code user-friendly wrappers. Claude Code best practices. Claude Code this, Claude Code that.&lt;/p&gt;

&lt;p&gt;It's reached that point in a tool's lifecycle where the cottage industry of tips and tricks is bigger than the tool's actual documentation. Which is either a sign of massive adoption or a sign that the thing is too fiddly to use without a guide. Probably both.&lt;/p&gt;

&lt;p&gt;The Claude Code Router bit is interesting though. It lets you intercept requests and route them to cheaper models for simple tasks. Which is exactly the kind of pragmatic cost-saving hack that makes the difference between "fun experiment" and "thing I can actually run in CI without my manager asking why the bill doubled."&lt;/p&gt;




&lt;h2&gt;
  
  
  DOGE Leaks an xAI API Key 🔑
&lt;/h2&gt;

&lt;p&gt;And because no week is complete without a government-adjacent security cock-up: DOGE accidentally exposed an xAI API key. In the clear. On the internet.&lt;/p&gt;

&lt;p&gt;I'm not going to pile on because frankly, &lt;em&gt;everyone&lt;/em&gt; has committed credentials at some point. But when you're a government entity with "efficiency" literally in the name, leaking API keys is a proper bad look. Especially when the key gives access to Grok, and the whole thing gets picked up by every security newsletter in existence.&lt;/p&gt;

&lt;p&gt;Here's what gets me though: we keep shipping AI tools faster than we ship the security practices to use them safely. API key management is a solved problem. Secret scanning is a solved problem. &lt;code&gt;git-secrets&lt;/code&gt;, &lt;code&gt;trufflehog&lt;/code&gt;, pre-commit hooks, all been around for years. And yet here we are, watching a government agency faff about with credentials like it's their first day on GitHub.&lt;/p&gt;




&lt;h2&gt;
  
  
  So Where Does This Leave Us?
&lt;/h2&gt;

&lt;p&gt;Zooming out from the wreckage of the week. Three companies fighting over one AI coding startup. China giving away trillion-parameter models like free samples at Costco. Newsletter hot takes about all models being the same. Every cloud provider launching their own IDE.&lt;/p&gt;

&lt;p&gt;We're watching AI coding become a commodity in real time. The models are converging, the features are all copying each other, and the pricing is in a death spiral towards zero.&lt;/p&gt;

&lt;p&gt;The winners won't be the ones with the best model. They'll be the ones with the best &lt;em&gt;workflow&lt;/em&gt;, the tightest integration between the model, the IDE, the CI pipeline, and the developer's actual habits. That's why Cursor is winning right now. Not because their model is better (they use other people's models). Because the experience of &lt;em&gt;using&lt;/em&gt; it is better.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; Vibe coding hit peak hype this week and I'm not sure the hype is wrong. The tools are getting proper good. But the messy acquisition drama, the security leaks, and the "just ship another IDE" energy all tell the same story: the AI coding market is moving faster than anyone's ability to make sense of it. If you missed &lt;a href="https://dev.to/byte-sized-banter/week-27-claude-vs-warp"&gt;the Claude Code vs Warp showdown&lt;/a&gt;, that tells you where the real differences lie. Strap in, crack on, and maybe rotate your API keys while you're at it.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>MCP Token Overhead: Why Model Context Protocol Is Kryptonite for AI Context Windows</title>
      <dc:creator>Steven Gonsalvez</dc:creator>
      <pubDate>Fri, 10 Jul 2026 15:42:37 +0000</pubDate>
      <link>https://dev.to/stevengonsalvez/mcp-token-overhead-why-model-context-protocol-is-kryptonite-for-ai-context-windows-3625</link>
      <guid>https://dev.to/stevengonsalvez/mcp-token-overhead-why-model-context-protocol-is-kryptonite-for-ai-context-windows-3625</guid>
      <description>&lt;h2&gt;
  
  
  Called It: MCP Security Is Already Falling Apart 🔓
&lt;/h2&gt;

&lt;p&gt;So that didn't take long.&lt;/p&gt;

&lt;p&gt;Oligo Security just dropped a &lt;a href="https://thehackernews.com/2025/07/critical-vulnerability-in-anthropics.html" rel="noopener noreferrer"&gt;critical RCE vulnerability in MCP Inspector&lt;/a&gt; - CVSS 9.4. The attack is beautifully simple: visit a dodgy website while running the inspector, attacker gets code execution on your machine. DNS rebinding chained with CSRF. Classic web attack vector, brand new attack surface.&lt;/p&gt;

&lt;p&gt;I wrote about &lt;a href="https://dev.to/blog/mcp-ai-security-risks-best-practices"&gt;exactly this back in May&lt;/a&gt;, covering tool poisoning, rug pulls, shadowing attacks, the full catalogue. Took about six weeks for the first real CVE to land. Not exactly chuffed about being right on this one, but the &lt;a href="https://www.securityweek.com/anthropic-mcp-server-flaws-lead-to-code-execution-data-exposure/" rel="noopener noreferrer"&gt;IDE extension exploits&lt;/a&gt; surfacing the same week just prove the point.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Bigger Problem: MCP Is Solving Something That Doesn't Need Solving 🧱
&lt;/h2&gt;

&lt;p&gt;The security holes are bad. But honestly? They're a symptom of a deeper problem.&lt;/p&gt;

&lt;p&gt;Here's my spicy take: &lt;strong&gt;MCP is the EAI bloat of the AI era.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We already have &lt;code&gt;curl&lt;/code&gt;. We have native APIs. We have CLIs. Every provider already exposes their functionality through well-documented, battle-tested interfaces. MCP is middleware nobody asked for, wedged between the provider and the consumer, bringing along an entirely new aggregated auth problem and unnecessary logic creeping into layers where it has no business being.&lt;/p&gt;

&lt;p&gt;If you've been around long enough to remember Enterprise Application Integration - the SOAP gateways, the ESBs, the message brokers that were supposed to unify everything and instead created a new class of problems - this should feel familiar. Same pitch. Same trajectory. Different decade.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;📚 Geek Corner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;The EAI parallel&lt;/strong&gt;: Enterprise Application Integration promised universal interoperability through middleware. What it delivered: XML hell, schema sprawl, vendor lock-in, and an entire industry of consultants debugging message transformations. MCP is walking the same path - a protocol layer between AI agents and tools that duplicates what HTTP APIs, CLIs, and native SDKs already provide, while introducing new failure modes and security surfaces.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Context Windows Are Gold. MCP Is Kryptonite. 💎
&lt;/h2&gt;

&lt;p&gt;This is the part that really gets me.&lt;/p&gt;

&lt;p&gt;The Playwright MCP server loads &lt;strong&gt;12.8k tokens&lt;/strong&gt; just by existing in your context window. Not &lt;em&gt;using&lt;/em&gt; it. Not running a single browser command. Just loading the tool definitions.&lt;/p&gt;

&lt;p&gt;Write a Python class skeleton with all the same browser automation functions? About &lt;strong&gt;1k tokens&lt;/strong&gt;. Dynamic scripts generated from that cost basically nothing in additional context.&lt;/p&gt;

&lt;p&gt;That's a 13x overhead for the privilege of running through a protocol layer.&lt;/p&gt;

&lt;p&gt;Context is going to be the gold of building good AI-assisted software. Every token in your context window is a trade-off - that's a token that could hold actual code, actual documentation, actual &lt;em&gt;reasoning&lt;/em&gt;. And MCP is hoovering up thousands of tokens on JSON tool schemas before you've even asked it to do anything.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# MCP Playwright: ~12,800 tokens just loaded
# Your tools: navigate, click, fill, screenshot, evaluate...
# All described in verbose JSON schemas
# None of them doing anything yet
&lt;/span&gt;
&lt;span class="c1"&gt;# Python wrapper: ~1,000 tokens total
&lt;/span&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Browser&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;navigate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="bp"&gt;...&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;selector&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="bp"&gt;...&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;fill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;selector&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;value&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="bp"&gt;...&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;screenshot&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;path&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="bp"&gt;...&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;evaluate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;js&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="bp"&gt;...&lt;/span&gt;
&lt;span class="c1"&gt;# Dynamic scripts from here cost nothing extra
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When context window size determines how much of your codebase an AI agent can reason about, burning 12k tokens on tool definitions is like paying rent on an office you never visit.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;📚 Geek Corner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;The context tax math&lt;/strong&gt;: MCP tool definitions get injected as JSON schemas - name, description, parameter types, examples, per tool. Playwright MCP has dozens of tools. At ~12.8k tokens, that's roughly 10% of a 128k context window gone before any work starts. In a 200k window, it's still 6.4%. Scale that across multiple MCP servers and you're losing a quarter of your reasoning capacity to protocol overhead. A lightweight native wrapper achieves the same functionality at 1/13th the context cost.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Where The Smart Money Is Going Instead 💰
&lt;/h2&gt;

&lt;p&gt;Ben Tossell ran a piece this week on &lt;a href="https://www.bensbites.com/p/im-investing-in-ai-devtools" rel="noopener noreferrer"&gt;investing in AI devtools&lt;/a&gt; - "bundling productivity is the new game."&lt;/p&gt;

&lt;p&gt;He's spot on. The standalone AI tool market is already getting crowded - everyone and their nan has a coding assistant. The winners will be the ones that bundle: code + tests + deployment + monitoring in one agentic flow. Not twenty MCP servers duct-taped together.&lt;/p&gt;

&lt;p&gt;The irony is thick: the MCP ecosystem is &lt;em&gt;unbundling&lt;/em&gt; by design - separate servers for everything, each burning context tokens, each with its own auth surface. Meanwhile the market is clearly moving toward &lt;em&gt;bundling&lt;/em&gt; - integrated tools that do more with less overhead.&lt;/p&gt;

&lt;p&gt;Whether that's Claude Code eating the terminal, Cursor eating the IDE, or something we haven't seen yet - the future is native integrations, not protocol middleware.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; Mark my words, MCP will go the way of SOAP, ESBs, and every middleware layer that promised interoperability and delivered complexity. The tools that win will be the ones that talk directly to what they need, not through yet another abstraction layer burning your context window for the privilege. For a deeper look at the architecture itself and where I reckon the real gaps are, see the &lt;a href="https://dev.to/blog/exploring-mcp-ecosystem-under-the-hood"&gt;MCP under the hood&lt;/a&gt; deep dive.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Claude Code vs Codex: iOS Simulator Build Test That Settled the Debate</title>
      <dc:creator>Steven Gonsalvez</dc:creator>
      <pubDate>Fri, 10 Jul 2026 15:41:58 +0000</pubDate>
      <link>https://dev.to/stevengonsalvez/claude-code-vs-codex-ios-simulator-build-test-that-settled-the-debate-41</link>
      <guid>https://dev.to/stevengonsalvez/claude-code-vs-codex-ios-simulator-build-test-that-settled-the-debate-41</guid>
      <description>&lt;h2&gt;
  
  
  Claude Code vs Codex: The iOS Simulator Test 📱
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Some tasks separate the AI coding agents from the autocompletes.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Building and running an iOS app in the simulator sounds simple. Open Xcode, hit the play button, done. But ask an AI coding agent to do it autonomously? That's where things get properly interesting.&lt;/p&gt;

&lt;p&gt;I threw the same challenge at OpenAI's Codex and Claude Code: "Build and run this app on the iOS simulator." Same machine, same permissions, same codebase. A proper head-to-head comparison of the two biggest AI coding tools on the market.&lt;/p&gt;




&lt;h3&gt;
  
  
  OpenAI Codex: The Graceful Surrender
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvy2ccd4q3f0l6x2evdkm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvy2ccd4q3f0l6x2evdkm.png" alt="Codex giving up on iOS simulator"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Codex tried. It really did. But after hitting the &lt;code&gt;CoreSimulatorService connection became invalid&lt;/code&gt; wall, it did something I actually respect: &lt;strong&gt;it gave up gracefully and handed me a manual&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;I tried to run it, but this environment can't access CoreSimulatorService,
so I can't launch the iOS Simulator from here. Your machine can, so here
are the exact local steps to deploy and run.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then it listed out the commands:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;npm run build&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;npx cap sync ios&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;npx cap run ios --target "iPhone 16"&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Fair enough. It recognised its limitations, explained &lt;em&gt;why&lt;/em&gt; it couldn't do the thing, and gave me a clear path forward. That's not failure - that's bounded rationality in action.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Feels like:&lt;/strong&gt; A contractor saying "I can't do electrical work, but here's exactly what you need to tell the electrician."&lt;/p&gt;




&lt;h3&gt;
  
  
  Claude Code: Just Does the Thing
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fstevengonsalvez.com%2Fimages%2Fbanter%2Fclaude-building-and-running-on-ios-simulator.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fstevengonsalvez.com%2Fimages%2Fbanter%2Fclaude-building-and-running-on-ios-simulator.png" alt="Claude Code running iOS simulator"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Claude Code, meanwhile, just... did it.&lt;/p&gt;

&lt;p&gt;The screenshot tells the story: checkboxes ticking off, build commands executing, simulator launching, app running. No drama. No apologies. No "here's a manual for you to do it yourself."&lt;/p&gt;

&lt;p&gt;It spawned subagents to research Phaser mobile scaling. It figured out the Capacitor config. It updated the native fetch settings. It ran the simulator. Done.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;📚 Geek Corner&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Same setup, different outcomes&lt;/strong&gt;: Both tools running locally, both with full access, both in YOLO mode. The difference is purely agentic capability - Claude Code's tooling and model combination just handles the complexity better. Codex hits a wall and falls back to "here's a manual." Claude Code hits the same wall and figures out how to get around it.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Claude Code's Subagent Architecture in Action 🤖
&lt;/h2&gt;

&lt;p&gt;Speaking of Claude doing the work, check out this little moment (this subagent pattern is part of what makes &lt;a href="https://stevengonsalvez.com/byte-sized-banter/week-33-terminal-wins" rel="noopener noreferrer"&gt;terminal agents so much more capable&lt;/a&gt; than IDE-based copilots):&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyu7kumh7uxc1hbkjl9cr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyu7kumh7uxc1hbkjl9cr.png" alt="Claude spawning an Explore subagent"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;"Let me use the &lt;strong&gt;Explore&lt;/strong&gt; agent to research how Phaser handles responsive design for mobile games with different aspect ratios"&lt;/p&gt;

&lt;p&gt;Then: &lt;strong&gt;Done (9 tool uses · 41.1k tokens · 31.9s)&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In half a minute, it spawned a research agent, burned through 41k tokens of documentation, and came back with an answer. "Excellent analysis! The agent confirmed that switching to RESIZE mode is the best solution."&lt;/p&gt;

&lt;p&gt;This is what agentic actually looks like. Not just answering questions, but &lt;em&gt;delegating research to itself&lt;/em&gt; and synthesising the results.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Feels like:&lt;/strong&gt; Asking a senior dev for help and watching them spin up a quick Slack thread with three other engineers, then coming back with "Sorted, here's the fix."&lt;/p&gt;




&lt;h2&gt;
  
  
  The Verdict
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Attempted iOS build?&lt;/th&gt;
&lt;th&gt;Actually ran?&lt;/th&gt;
&lt;th&gt;Handled gracefully?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Code&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;N/A - just worked&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Codex&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes - gave manual&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Same machine. Same access. Same YOLO permissions. Different results.&lt;/p&gt;

&lt;p&gt;Codex encountered friction and decided "this is too hard, here's a manual." Claude Code encountered the same friction and &lt;em&gt;worked through it&lt;/em&gt;. That's the difference between autocomplete-with-attitude and actual agentic problem-solving.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; When both AI coding agents have the same access and one just gives up while the other figures it out, that's not a constraint issue. That's a capability gap.&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
