<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alvarito1983</title>
    <description>The latest articles on DEV Community by Alvarito1983 (@alvarito1983).</description>
    <link>https://dev.to/alvarito1983</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3858088%2Fa0fbc217-d69f-4570-b98d-c85a5e01ed1d.png</url>
      <title>DEV Community: Alvarito1983</title>
      <link>https://dev.to/alvarito1983</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/alvarito1983"/>
    <language>en</language>
    <item>
      <title>GTA VI's Extended Look didn't just show gameplay — it answered the leak controversy without saying a word</title>
      <dc:creator>Alvarito1983</dc:creator>
      <pubDate>Fri, 28 Aug 2026 07:51:48 +0000</pubDate>
      <link>https://dev.to/alvarito1983/gta-vis-extended-look-didnt-just-show-gameplay-it-answered-the-leak-controversy-without-saying-1k9j</link>
      <guid>https://dev.to/alvarito1983/gta-vis-extended-look-didnt-just-show-gameplay-it-answered-the-leak-controversy-without-saying-1k9j</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published in Spanish on &lt;a href="https://elrack.es/gaming/gta-6-extended-look-netflix-analisis/" rel="noopener noreferrer"&gt;El Rack&lt;/a&gt;. Browser translation handles the rest of the site fine if you're into homelab/security content.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A few days ago I wrote about &lt;a href="https://elrack.es/seguridad/filtracion-hackeo-gta-6-cyberleek-2026/" rel="noopener noreferrer"&gt;the Cyberleek leak&lt;/a&gt; and its manifesto demanding Rockstar ship GTA VI on physical disc. On August 27, Rockstar answered — just not in words. "Grand Theft Auto VI: An Extended Look" premiered on Netflix with a 6-hour timed exclusive, 26 minutes of footage captured entirely on PS5, and one detail that probably isn't a coincidence of timing: physical copies won't include a disc, only a download code. The exact opposite of what the group behind the breach demanded.&lt;/p&gt;

&lt;h2&gt;
  
  
  The response nobody expected (or maybe everybody did)
&lt;/h2&gt;

&lt;p&gt;There's no official statement from Rockstar linking the two, and there probably never will be — caving to an extortion threat would set an awful precedent for any company. But the sequence is hard to ignore: a manifesto demanding physical media, and days later, official confirmation that even the physical editions ship without a disc. Intentional or not, the message anyone paying attention receives is clear: pressure from a hacker group doesn't move commercial decisions that were already made.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Caving would have been the story. Not caving, instead, barely registers — and that silence says more than any statement could.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The price that resets the industry's ceiling
&lt;/h2&gt;

&lt;p&gt;GTA VI's $79.99 marks a real jump from the $69.99 that had already become standard this generation — and with development costs reportedly well north of $200 million, the move makes financial sense even if nobody's going to like it. What matters isn't just the number itself, it's what it probably unlocks: if the most anticipated release of the decade sets $80 as the reference price, it's a matter of time before other major launches follow, the same way the industry jumped from $59.99 to $69.99 a few years back.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gameplay details that actually matter, past the hype
&lt;/h2&gt;

&lt;p&gt;Buried under headlines about runtime and pricing, there are mechanic confirmations that deserve more attention than they're getting. The footage confirms a &lt;strong&gt;karma system&lt;/strong&gt; — petting a dog, for instance, earns points, in the same vein as Red Dead Redemption 2 — that will shape how the world reacts to your actions, something rumors had been pointing at for months without official confirmation until now. There's also a &lt;strong&gt;team cover system&lt;/strong&gt; during shootouts, real-time character switching between Jason and Lucia, and vehicles built specifically for Leonida's wilder terrain — swamp sliders, kayaks, electric scooters — though it's unclear yet whether the last one is player-drivable or NPC-only.&lt;/p&gt;

&lt;p&gt;Confirmed in-game Spanish dialogue, plus Rockstar's own explicit references to classic action cinema (Bad Boys, Heat) as narrative influence, point to writing ambitions beyond the series' usual formula.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's still unresolved
&lt;/h2&gt;

&lt;p&gt;With no PC version confirmed at launch — Rockstar's usual pattern, GTA V took nearly a year to reach PC after consoles — anyone who plays primarily on PC will be waiting well past November 19. And while the footage clears up plenty of mechanics questions, things like online mode or post-launch content still have zero official mention.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;Past the spectacle (26 minutes, Netflix exclusive, half a billion cumulative trailer views), the most interesting part of this extended look is what it confirms without saying it: neither leak pressure nor a hacker group's manifesto moved the commercial direction of a release at this scale. The game, unplayed as it still is, looks set to deliver on the technical and mechanical promises it's been making for years — the price and the disc-less physical editions, on the other hand, are the part of this story that's going to keep generating conversation.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Full article (in Spanish): &lt;a href="https://elrack.es/gaming/gta-6-extended-look-netflix-analisis/" rel="noopener noreferrer"&gt;elrack.es/gaming/gta-6-extended-look-netflix-analisis&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>gaming</category>
      <category>news</category>
      <category>discuss</category>
    </item>
    <item>
      <title>The GTA VI leak isn't really about GTA VI — it's an extortion playbook</title>
      <dc:creator>Alvarito1983</dc:creator>
      <pubDate>Thu, 20 Aug 2026 11:41:23 +0000</pubDate>
      <link>https://dev.to/alvarito1983/the-gta-vi-leak-isnt-really-about-gta-vi-its-an-extortion-playbook-o1a</link>
      <guid>https://dev.to/alvarito1983/the-gta-vi-leak-isnt-really-about-gta-vi-its-an-extortion-playbook-o1a</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published in Spanish on &lt;a href="https://elrack.es/gaming/filtracion-hackeo-gta-6-cyberleek-2026/" rel="noopener noreferrer"&gt;El Rack&lt;/a&gt;. Browser translation handles the rest of the site fine if you're into homelab/security content.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;On August 18, 2026, nine days before GTA VI's official Netflix reveal, two videos claiming to show gameplay footage started circulating, along with an image of what's supposedly the full Vice City map. A group calling itself Cyberleek claimed responsibility, and what started as another leak has turned into something with the shape of a real security incident — manifesto, demands, and threats of more material if those demands aren't met.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's confirmed, and what isn't
&lt;/h2&gt;

&lt;p&gt;The footage shows Jason, one of the game's protagonists, at the beach house already seen in the second official trailer, shooting hoops in the yard — with technical details (a coin-pickup-style sound effect on scoring) the community is reading as an authenticity signal. Growing DMCA takedown activity on the shared content reinforces that reading, though as with any leak, the possibility of AI-generated footage means caution is warranted until there's official confirmation. Some reports suggest the material could be over a year old, which would fit with a leak from an older development build rather than the game's current state.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"If we can reach Rockstar, no one is safe" isn't the language of a casual leak — it's extortion-manifesto language, the same pattern we've seen in corporate attacks well outside the gaming industry.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The manifesto: this isn't just about GTA VI
&lt;/h2&gt;

&lt;p&gt;Cyberleek frames the attack as protest against Rockstar and Take-Two's decision not to ship a physical disc release, and the shared text goes further, attacking the industry's broader monetization playbook: games-as-a-service, digital licenses instead of real ownership, disc-locked content behind paid DLC. The group demands publishers "behave" and explicitly threatens to extend attacks to other companies if demands aren't met — a message aimed at the whole industry, not just Rockstar.&lt;/p&gt;

&lt;p&gt;There's a contradiction worth flagging: the manifesto insists "this isn't for money," but parallel reports suggest the group is charging for a poll to pick the next leak — the classic double-talk of modern extortion, stated principles in public, monetization in private.&lt;/p&gt;

&lt;h2&gt;
  
  
  How this compares to the 2022 precedent
&lt;/h2&gt;

&lt;p&gt;In September 2022, an attacker identifying as "teapotuberhacker" (linked to the Lapsus$ collective) published roughly 90 videos from a very early development build, in what Rockstar officially confirmed as a real intrusion into their servers — not a careless employee leak, but an actual technical compromise of their infrastructure. The current leak, per the outlets covering it, is the second-largest since then, though visibly smaller in scope than 2022 (2 videos vs. 90).&lt;/p&gt;

&lt;h2&gt;
  
  
  The context that makes this more flammable
&lt;/h2&gt;

&lt;p&gt;GTA VI has already taken two consecutive delays — from fall 2025 to May 2026, then to November 2026 — with Take-Two's own CEO acknowledging the current timeline runs roughly 18 months behind the original plan. With a fanbase running on years of anticipation, any leak, real or not, finds fertile ground to go viral before the company can respond officially.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;Beyond the gaming angle, this is a textbook case study in modern digital extortion: a real intrusion (or leak), a public manifesto with explicit threats against third parties, and covert monetization despite the "not about money" framing. Worth tracking for anyone in security — less for GTA VI itself, more for what it reveals about how these groups operate today.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Full article (in Spanish): &lt;a href="https://elrack.es/gaming/filtracion-hackeo-gta-6-cyberleek-2026/" rel="noopener noreferrer"&gt;elrack.es/seguridad/filtracion-hackeo-gta-6-cyberleek-2026&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>gaming</category>
      <category>infosec</category>
      <category>news</category>
    </item>
    <item>
      <title>MCP, Subagents, and Hooks in Claude Code: The Guide I Wish I'd Had</title>
      <dc:creator>Alvarito1983</dc:creator>
      <pubDate>Wed, 19 Aug 2026 08:36:15 +0000</pubDate>
      <link>https://dev.to/alvarito1983/mcp-subagents-and-hooks-in-claude-code-the-guide-i-wish-id-had-4gg</link>
      <guid>https://dev.to/alvarito1983/mcp-subagents-and-hooks-in-claude-code-the-guide-i-wish-id-had-4gg</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://elrack.es" rel="noopener noreferrer"&gt;El Rack&lt;/a&gt; — Spanish tech reviews from a sysadmin/homelab perspective.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Claude Code is capable right out of the box, but it starts to feel thin the moment your project needs to touch external systems: a ticket tracker, a database, your own VPS. This guide covers the four pieces that turn it into a real production tool:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;MCP&lt;/strong&gt; (&lt;code&gt;/mcp&lt;/code&gt;): connecting external services as tools&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Subagents&lt;/strong&gt;: delegating tasks to isolated instances with their own context&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skills / slash commands&lt;/strong&gt;: packaging repetitive prompts as commands&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hooks&lt;/strong&gt;: deterministic guardrails that don't depend on the AI "remembering" an instruction&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You'll need Claude Code installed and authenticated, and a terminal open in any project folder.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Connect your first MCP server
&lt;/h2&gt;

&lt;p&gt;MCP (Model Context Protocol) lets Claude Code use tools it doesn't ship with by default. Those tools live in MCP servers: local processes or hosted services reachable over a URL. You add them with &lt;code&gt;claude mcp add&lt;/code&gt;, no hand-editing JSON required. Start with the official docs server — it needs no account and no config:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add &lt;span class="nt"&gt;--transport&lt;/span&gt; http claude-code-docs https://code.claude.com/docs/mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Confirm it connected:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp list
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You should see &lt;code&gt;✔ Connected&lt;/code&gt;. From inside a session, manage servers any time with &lt;code&gt;/mcp&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Local (stdio) servers and OAuth-gated servers
&lt;/h2&gt;

&lt;p&gt;A stdio server is a program Claude Code launches as a subprocess — useful when it needs access to your filesystem or a browser. Example with Playwright, which requires no account:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add playwright &lt;span class="nt"&gt;--&lt;/span&gt; npx &lt;span class="nt"&gt;-y&lt;/span&gt; @playwright/mcp@latest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;--&lt;/code&gt; separates Claude Code's own flags from the command that starts the server. For services that require sign-in (Sentry, Linear, Notion, GitHub), you add them the same way and authenticate from inside the session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add &lt;span class="nt"&gt;--transport&lt;/span&gt; http sentry https://mcp.sentry.dev/mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After adding, you'll see &lt;code&gt;! Needs authentication&lt;/code&gt;. Start a session, run &lt;code&gt;/mcp&lt;/code&gt;, select the server, and choose "Authenticate" — your browser opens for sign-in.&lt;/p&gt;

&lt;p&gt;If the service uses a static token instead of OAuth (common with self-hosted instances), pass it directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude mcp add &lt;span class="nt"&gt;--transport&lt;/span&gt; http my-server http://my-host:3001/api/mcp &lt;span class="nt"&gt;--header&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer YOUR_TOKEN"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  3. Pick the right scope: local, project, or user
&lt;/h2&gt;

&lt;p&gt;By default, every server is registered at "local" scope: private to you, active only in the current project. Two alternatives depending on how widely you want to share it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Available in all your projects, still private&lt;/span&gt;
claude mcp add &lt;span class="nt"&gt;--scope&lt;/span&gt; user &lt;span class="nt"&gt;--transport&lt;/span&gt; http claude-code-docs https://code.claude.com/docs/mcp

&lt;span class="c"&gt;# Shared with the team via the repo (writes .mcp.json)&lt;/span&gt;
claude mcp add &lt;span class="nt"&gt;--scope&lt;/span&gt; project &lt;span class="nt"&gt;--transport&lt;/span&gt; http claude-code-docs https://code.claude.com/docs/mcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;Available to&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;local&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;~/.claude.json&lt;/code&gt; (project entry)&lt;/td&gt;
&lt;td&gt;Only you, only this project&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;project&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;.mcp.json&lt;/code&gt; at the repo root&lt;/td&gt;
&lt;td&gt;Everyone who clones the repo&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;user&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;~/.claude.json&lt;/code&gt; (top-level &lt;code&gt;mcpServers&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Only you, all your projects&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you'd rather write &lt;code&gt;.mcp.json&lt;/code&gt; by hand to keep it version-controlled with the team:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"claude-code-docs"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://code.claude.com/docs/mcp"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"playwright"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"stdio"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"npx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"args"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"-y"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"@playwright/mcp@latest"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  4. Create your first custom subagent
&lt;/h2&gt;

&lt;p&gt;A subagent is an instance with its own context, its own system prompt, and its own tools, working in isolation and returning only a summary. It's the fix for two problems: flooding your main conversation with logs or results you won't reuse, and re-spawning the same kind of worker with the same instructions over and over.&lt;/p&gt;

&lt;p&gt;Claude Code ships with three built-in subagents: Explore (read-only code search), Plan (research during plan mode), and general-purpose (complex tasks with access to everything). For a custom one, just ask Claude to write it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Create a personal subagent in ~/.claude/agents/ called "code-reviewer" that
reviews code for quality, security, and best practices. Make it read-only
and have it use Sonnet.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Claude writes the file with YAML frontmatter plus the system prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;code-reviewer&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Reviews code for quality, security, and best practices. Use after writing or modifying code.&lt;/span&gt;
&lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Read, Grep, Glob&lt;/span&gt;
&lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sonnet&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="s"&gt;You are a senior code reviewer. For each issue you find, explain the&lt;/span&gt;
&lt;span class="s"&gt;problem, show the current code, and provide an improved version.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The most useful frontmatter fields: &lt;code&gt;tools&lt;/code&gt; (allowlist), &lt;code&gt;disallowedTools&lt;/code&gt; (denylist), &lt;code&gt;model&lt;/code&gt; (sonnet/opus/haiku/inherit), and &lt;code&gt;mcpServers&lt;/code&gt; (gives an MCP server to that subagent alone, without loading its context into the main conversation).&lt;/p&gt;

&lt;p&gt;Save it in &lt;code&gt;.claude/agents/&lt;/code&gt; to scope it to this project, or &lt;code&gt;~/.claude/agents/&lt;/code&gt; to make it available everywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Invoking subagents: natural language, @-mention, or a full session
&lt;/h2&gt;

&lt;p&gt;There are three ways to use a subagent, from least to most explicit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Natural language: Claude decides whether to delegate
Use the code-reviewer subagent to review my recent changes

# @-mention: forces that specific subagent to run
@"code-reviewer (agent)" review the authentication logic

# Run the whole session as that subagent (its system prompt and tools)
claude --agent code-reviewer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For independent investigations, you can request several subagents in parallel:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Research the authentication, database, and API modules in parallel using
separate subagents
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each one explores its own area in isolation; you only get the synthesized summary back in your main conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Package repetitive prompts: slash commands and skills
&lt;/h2&gt;

&lt;p&gt;If you type the same instruction over and over, turn it into a command. Classic commands (&lt;code&gt;.claude/commands/*.md&lt;/code&gt;) still work, but the recommended approach now is Skills (&lt;code&gt;.claude/skills/&amp;lt;name&amp;gt;/SKILL.md&lt;/code&gt;) — if a command and a skill share a name, the skill wins.&lt;/p&gt;

&lt;p&gt;Classic command, saved as &lt;code&gt;.claude/commands/audit-disk.md&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Disk space audit with a configurable threshold&lt;/span&gt;
&lt;span class="na"&gt;allowed-tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Read, Bash, Grep&lt;/span&gt;
&lt;span class="na"&gt;argument-hint&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;threshold-percentage&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="na"&gt;Audit disk space across the servers. Alert threshold&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;$ARGUMENTS%.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Invoke it with &lt;code&gt;/audit-disk 85&lt;/code&gt;. &lt;code&gt;$ARGUMENTS&lt;/code&gt; captures everything typed after the command; &lt;code&gt;!`command`&lt;/code&gt; injects live shell output (e.g. &lt;code&gt;!`git diff --cached`&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;Same idea as a skill, at &lt;code&gt;.claude/skills/audit-disk/SKILL.md&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;audit-disk&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Disk space audit with a configurable threshold. Use when the user asks to check disk space on the servers.&lt;/span&gt;
&lt;span class="na"&gt;allowed-tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Read, Bash, Grep&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="s"&gt;Audit disk space across the servers...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The advantage of skills: they can bundle several reference files in the same folder, and Claude can invoke them on its own, without you typing the slash.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. Put deterministic guardrails in place with hooks
&lt;/h2&gt;

&lt;p&gt;If you run in an autonomous mode (&lt;code&gt;--dangerously-skip-permissions&lt;/code&gt;), hooks are how you block dangerous operations without relying on the AI remembering a rule. They fire every time, at the exact lifecycle point you define.&lt;/p&gt;

&lt;p&gt;The two most-used events: &lt;code&gt;PreToolUse&lt;/code&gt; (before a tool runs — good for blocking) and &lt;code&gt;PostToolUse&lt;/code&gt; (after — good for formatting, linting, or logging). Configure them in &lt;code&gt;settings.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"PreToolUse"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"matcher"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Bash"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"./scripts/validate-command.sh"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The script receives the command as JSON via stdin, and exit code 2 blocks the operation. Example that stops an overly broad &lt;code&gt;pkill&lt;/code&gt; on a server where multiple processes share the same name:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/bin/bash&lt;/span&gt;
&lt;span class="nv"&gt;INPUT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;COMMAND&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$INPUT&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | jq &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s1"&gt;'.tool_input.command // empty'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$COMMAND&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-qE&lt;/span&gt; &lt;span class="s1"&gt;'pkill.*node|rm -rf /'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
  &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Blocked: command too broad or destructive. Use an exact PID."&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&amp;amp;2
  &lt;span class="nb"&gt;exit &lt;/span&gt;2
&lt;span class="k"&gt;fi

&lt;/span&gt;&lt;span class="nb"&gt;exit &lt;/span&gt;0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Make it executable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;chmod&lt;/span&gt; +x ./scripts/validate-command.sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Another common one: auto-format every file Claude touches, without having to ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"PostToolUse"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"matcher"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Write|Edit"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"npx prettier --write &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;$CLAUDE_TOOL_INPUT_FILE_PATH&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Wrapping up
&lt;/h2&gt;

&lt;p&gt;Put these four pieces together and Claude Code stops being a terminal assistant and becomes a real automation layer: it connects to whatever it needs (MCP), delegates whatever would clutter its context (subagents), packages whatever repeats (skills), and respects hard limits that don't depend on its memory (hooks). Suggested adoption order: start with one low-risk MCP server, add one read-only subagent, migrate your commands to skills whenever you have time, and — the highest-leverage one if you run in autonomous mode — add at least one hook that blocks the operation you're most afraid of running by accident.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Tutorial by Álvaro Fraguas Bravo for &lt;a href="https://elrack.es/tutoriales/herramientas-ia/claude-code-mcp-subagentes-hooks/" rel="noopener noreferrer"&gt;El Rack&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claudecode</category>
      <category>mcp</category>
      <category>automation</category>
      <category>devops</category>
    </item>
    <item>
      <title>Mistral Shieldstral 1.0 Review — A 3B Self-Hostable Moderation Model That Runs on a Single 16GB GPU</title>
      <dc:creator>Alvarito1983</dc:creator>
      <pubDate>Wed, 19 Aug 2026 07:02:10 +0000</pubDate>
      <link>https://dev.to/alvarito1983/mistral-shieldstral-10-review-a-3b-self-hostable-moderation-model-that-runs-on-a-single-16gb-gpu-3ecb</link>
      <guid>https://dev.to/alvarito1983/mistral-shieldstral-10-review-a-3b-self-hostable-moderation-model-that-runs-on-a-single-16gb-gpu-3ecb</guid>
      <description>&lt;p&gt;On August 5, 2026, Mistral released &lt;strong&gt;Shieldstral 1.0&lt;/strong&gt;, a 3-billion-parameter model built on top of Ministral-3-3B-Base-2512, designed to do exactly one job: moderate text and image content before it reaches an end user. What makes it interesting for a homelab audience isn't just what it does, but how it's shipped — full weights on Hugging Face under an Apache 2.0 license, with no fine print around self-hosting.&lt;/p&gt;

&lt;h2&gt;
  
  
  A model that fits the hardware you already own
&lt;/h2&gt;

&lt;p&gt;Unlike the hundred-billion-plus-parameter frontier models we usually cover in this section, Shieldstral is built to fit on a &lt;strong&gt;single 16GB GPU in BF16 precision&lt;/strong&gt; — a card a good chunk of the local-AI homelab crowd already owns, or can justify without selling a kidney. It supports the usual homelab inference stack: vLLM, llama.cpp, SGLang, and Transformers, plus Axolotl if you want to fine-tune it on your own policy set.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What sets Shieldstral apart isn't that it moderates well — it's that it does so without forcing you to send your content to someone else's moderation API.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The "policy-adaptive" feature is the most practical piece of this release: instead of retraining the model every time a community's, forum's, or project's rules change, the policy is described in natural language directly in the prompt. It's the same principle other recent specialized models have adopted, applied here to a very specific use case — multimodal moderation, not general-purpose text generation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmarks Mistral published, not yet independently verified
&lt;/h2&gt;

&lt;p&gt;The numbers Mistral reports look strong on paper: 99.4% F1 on HarmBench, 97.7% on the multimodal VLGuard set, 84.1% on ToxicChat. These are vendor-reported figures — as of this review, there's no independent comparison benchmarking them against other self-hosted or commercial moderation solutions. That's not a reason to dismiss the model, but it is a reason not to take them at face value if you're building something critical on top of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this makes sense for right now
&lt;/h2&gt;

&lt;p&gt;If you're already running your own service — a forum, a community, a platform with user-generated content — and you want moderation without depending on a third party's API or sending sensitive content outside your network, Shieldstral is a real option with auditable weights worth trying. If you need a general-purpose model for chat or agents, this isn't it — it's a specialized piece meant to sit alongside another model, not replace it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Specs at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Release date&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;August 5, 2026&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Parameters&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;3 billion, on top of the Ministral-3-3B-Base-2512 base model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;License&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Apache 2.0 — full weights downloadable on Hugging Face (&lt;code&gt;mistralai/Shieldstral-1.0-3B&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Hardware requirement&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Fits on a single 16GB VRAM GPU in BF16 precision&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Supported inference frameworks&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;vLLM (0.26.0+), llama.cpp, SGLang, Transformers, and Axolotl for fine-tuning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Supported languages&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;12 — English, French, Spanish, German, Italian, Portuguese, Dutch, Chinese, Japanese, Korean, Arabic, and Russian&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Context&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Trained up to 32,000 tokens, with theoretical support up to 256,000 tokens (unvalidated by Mistral's own benchmarks at that length)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Benchmarks (F1)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;HarmBench 99.4% — ToxicChat 84.1% — VLGuard (multimodal) 97.7% — XSTest refusal 94.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  How it compares
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Seed 2.1 Turbo (ByteDance)&lt;/strong&gt; — the direct contrast in the opposite direction: decent agentic capabilities, but closed from the ground up, with no weights and no self-hosting option.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek V4-Pro&lt;/strong&gt; — another recent open-weight model under an MIT license, though general-purpose. Shieldstral instead bets on solving a single problem (moderation) rather than competing across the board.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Pros
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Full weights downloadable under Apache 2.0 — can be self-hosted, audited, and modified without depending on a third party's moderation API&lt;/li&gt;
&lt;li&gt;Fits on a single 16GB GPU in BF16, well within the hardware range a good part of the local-AI community we cover on El Rack already has&lt;/li&gt;
&lt;li&gt;Solid benchmarks on the toughest tests, with 99.4% F1 on HarmBench and 97.7% on the multimodal VLGuard set&lt;/li&gt;
&lt;li&gt;"Policy-adaptive" moderation lets you describe policy in natural language directly in the prompt, without retraining the model every time a community's or project's rules change&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Cons
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;It's a specialized yes/no classifier for content, not a general-purpose model — you'll need a separate model for the rest of your AI stack&lt;/li&gt;
&lt;li&gt;Published benchmarks are self-reported by Mistral — there's no independent comparison yet against other self-hosted moderation solutions&lt;/li&gt;
&lt;li&gt;Real training context is 32,000 tokens — the "theoretical" support up to 256,000 tokens isn't backed by any benchmark at that length&lt;/li&gt;
&lt;li&gt;12 languages is decent coverage, but not complete — it leaves out entire communities that some closed commercial solutions do serve&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Verdict — 7.6/10
&lt;/h2&gt;

&lt;p&gt;A specialized model with fully open weights, built to fit on hardware a lot of the local-AI homelab already has — the most serious self-hosted moderation offering we've seen lately. The benchmarks are good, but they come only from the vendor itself, and how useful it is to you depends entirely on whether you need this exact use case: moderation, not a general-purpose model.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Review by Lucía Fernández Moreno for &lt;a href="https://elrack.es/herramientas-ia/mistral-shieldstral-1-0-moderacion-local/" rel="noopener noreferrer"&gt;El Rack&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
      <category>opensource</category>
    </item>
    <item>
      <title>GitHub Copilot Agent Plugins 1.0 Goes GA — A Real Plugin System, Not Just a Roadmap Promise</title>
      <dc:creator>Alvarito1983</dc:creator>
      <pubDate>Mon, 17 Aug 2026 08:27:25 +0000</pubDate>
      <link>https://dev.to/alvarito1983/github-copilot-agent-plugins-10-goes-ga-a-real-plugin-system-not-just-a-roadmap-promise-4hck</link>
      <guid>https://dev.to/alvarito1983/github-copilot-agent-plugins-10-goes-ga-a-real-plugin-system-not-just-a-roadmap-promise-4hck</guid>
      <description>&lt;p&gt;On August 12, GitHub confirmed general availability for Agent Plugins 1.0, an extensibility system for Copilot's agent mode that had been in preview for a few weeks (the first note dates back to August 6). This isn't a single-client feature: the same plugin system works across VS Code, the Copilot CLI, the SDK, and the Copilot app itself, and it's available from day one on every paid subscription tier.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real extensibility, not a roadmap promise
&lt;/h2&gt;

&lt;p&gt;The most notable part of this release is that it ships with plugin management built for a real ecosystem, not a demo: every installed plugin gets its own version tracking, plus a one-click "update all" button. It's the kind of maintenance detail you only miss once you already have several plugins installed and one of them breaks something after a silent update — GitHub got ahead of that problem at launch instead of bolting it on later as a patch.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question worth asking before installing your first third-party plugin
&lt;/h2&gt;

&lt;p&gt;A plugin system that runs inside Copilot's agent mode — with potential access to your code and to the actions the agent can take on your behalf — opens up a new attack surface by definition, and GitHub hasn't yet publicly detailed what review process or sandboxing applies to a third-party plugin before it reaches users.&lt;/p&gt;

&lt;p&gt;That's not a reason to write off the feature — it's exactly the same question worth asking about any new plugin ecosystem, from VS Code to browsers — but it does mean installing cautiously until it's clearer what guarantees actually sit behind the system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Part of a broader move toward openness
&lt;/h2&gt;

&lt;p&gt;This update lands in the same window where Copilot added Kimi K3 and MAI-Code-1.1-Flash as additional selectable models — two third-party models with no ties to OpenAI. Read alongside Agent Plugins 1.0, the pattern is clear: GitHub is opening up Copilot in more than one direction at once — third-party models on one side, third-party plugins on the other — instead of keeping it a closed box with a single model provider and no extension path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who this makes sense for right now
&lt;/h2&gt;

&lt;p&gt;If you already pay for Copilot and want to extend agent mode beyond what ships out of the box, it's worth trying — version tracking and update management are better thought out than usual for a 1.0 launch. If your priority is privacy and full control over what code runs inside your agent workflow, wait until GitHub clarifies the plugin review process before installing anything from an author you don't already trust.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;GA date&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;August 12, 2026 (preview note on August 6), per GitHub's official changelog&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;What it is&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;a plugin system for GitHub Copilot's agent mode — extends the agent runtime, not a single-editor feature&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Platforms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;VS Code, GitHub Copilot CLI, GitHub Copilot SDK, and the Copilot app&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Plans&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;available on every Copilot tier: Pro, Pro+, Business, and Enterprise&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Installed-plugin management&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;per-plugin version tracking, with one-click "update all"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Same-week context&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Copilot also added Kimi K3 and MAI-Code-1.1-Flash as selectable models in the same period&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Verdict — 7.1/10
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A well-built extensibility system from launch, with plugin management designed for a real ecosystem and available on every paid tier from day one — but GitHub still hasn't clarified what review or sandboxing applies to a third-party plugin, which is the question to ask before letting any outside code run inside your agent workflow.&lt;/p&gt;

&lt;p&gt;Originally published in Spanish on &lt;a href="https://elrack.es/software/github-copilot-agent-plugins-1-0/" rel="noopener noreferrer"&gt;ElRack.es&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>github</category>
      <category>software</category>
      <category>softwaredevelopment</category>
    </item>
    <item>
      <title>Meta's Muse Glimmer: a real agentic model that fits on your own GPU</title>
      <dc:creator>Alvarito1983</dc:creator>
      <pubDate>Tue, 11 Aug 2026 08:58:02 +0000</pubDate>
      <link>https://dev.to/alvarito1983/metas-muse-glimmer-a-real-agentic-model-that-fits-on-your-own-gpu-16oh</link>
      <guid>https://dev.to/alvarito1983/metas-muse-glimmer-a-real-agentic-model-that-fits-on-your-own-gpu-16oh</guid>
      <description>&lt;p&gt;On August 10th, Meta released Muse Glimmer, a 30-billion-parameter open model built specifically to run local AI agents on consumer hardware — scheduling, file organization, tool calling, multi-step tasks — without shipping any of it to Meta's servers. Weights are on Hugging Face under an Apache 2.0 license, so there's no commercial restriction on using it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this one is different from the last wave of "local" models
&lt;/h2&gt;

&lt;p&gt;The 7B and 13B models that filled up Ollama's library over the last couple of years were solid text generators, but weak agents. Multi-step reasoning fell apart, tool calls failed halfway through, and they'd lose track of what happened three steps back in a task.&lt;/p&gt;

&lt;p&gt;Glimmer is Meta's answer to that gap specifically — it's a simplified, efficiency-focused derivative of Meta's larger Muse Spark 1.2 model, purpose-built for "always-on" agentic workflows: the kind of thing that needs to keep running continuously on a personal machine rather than answering one prompt at a time.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it actually needs to run
&lt;/h2&gt;

&lt;p&gt;30B parameters, quantized memory footprint of 18–20 GB — within reach of current high-end consumer GPUs, no server rack required&lt;br&gt;
Handles interleaved text and images (screenshots, documents, mixed content) across 100+ languages&lt;br&gt;
Native support for the frameworks local-AI users already have installed: Ollama, LM Studio, llama.cpp, MLX, vLLM, SGLang — if you already pull models from Hugging Face, this drops into your existing workflow with no extra tooling&lt;br&gt;
Meta is also working with AMD, Arm, Dell, Intel, and NVIDIA on hardware-specific performance tuning, and providing documentation for building custom agent scaffolds on top of it&lt;br&gt;
The bigger picture&lt;/p&gt;

&lt;p&gt;Glimmer didn't launch alone — Meta paired it with a 14-page essay from Mark Zuckerberg ("The Future Is for Everyone") arguing against AI capability staying locked inside a handful of companies, and confirmed weights for the larger Muse Spark model are coming too. It's a direct response to the pressure open models are putting on the market: Chinese labs — Moonshot's Kimi K3, Alibaba's Qwen3.8-Max, DeepSeek's V4-Flash — are shipping performance that rivals closed US frontier labs, and open weights are consistently cheaper to run at scale, which matters to anyone watching their own inference bill.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this fits in a self-hosted setup
&lt;/h2&gt;

&lt;p&gt;For anyone already running local models as part of a homelab — Ollama containers, a GPU passed through to a VM, that kind of setup — Glimmer is one of the first local models actually aimed at running unattended agent workflows instead of just answering chat prompts. That's a meaningfully different use case from "local chatbot," and worth testing against real tasks rather than benchmarks.&lt;/p&gt;

&lt;p&gt;I run local AI as part of my own homelab, so I'll be testing Glimmer against real agentic tasks rather than benchmarks — full writeup on El Rack once I've put it through its paces:&lt;/p&gt;

&lt;p&gt;👉 &lt;a href="https://elrack.es/herramientas-ia/muse-glimmer-meta/" rel="noopener noreferrer"&gt;Muse Glimmer, el nuevo modelo de Meta&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;(Spanish-language site — translation tools handle it cleanly if you don't read Spanish.)&lt;/p&gt;

&lt;p&gt;I write about homelab, self-hosting, and local AI at El Rack — real testing from inside my own homelab, not just a spec sheet rewrite.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
    </item>
    <item>
      <title>NVIDIA 610.57.04 for Linux: the point release actually worth installing</title>
      <dc:creator>Alvarito1983</dc:creator>
      <pubDate>Mon, 10 Aug 2026 12:03:11 +0000</pubDate>
      <link>https://dev.to/alvarito1983/nvidia-6105704-for-linux-the-point-release-actually-worth-installing-2phf</link>
      <guid>https://dev.to/alvarito1983/nvidia-6105704-for-linux-the-point-release-actually-worth-installing-2phf</guid>
      <description>&lt;p&gt;On August 3rd, NVIDIA quietly shipped two Linux driver branches on the same day: the new feature branch 610.57.04 and the stable 595.91.07. The stable release is the usual "minor bug fixes and improvements" non-answer, but 610.57.04 is a different story — it carries an unusually large fix list for a point release, and if you're gaming on Linux with an NVIDIA card, it's worth grabbing now instead of leaving it to rot in your package manager.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually got fixed
&lt;/h2&gt;

&lt;p&gt;More than 20 titles are patched in this release, several of which had been broken for weeks or months:&lt;/p&gt;

&lt;p&gt;STALKER 2: Heart of Chornobyl — crash with VK_DEVICE_LOST&lt;br&gt;
Assetto Corsa EVO — incorrectly wrapped textures&lt;br&gt;
Monster Hunter Wilds — crash at the main menu (driver reporting the GPU as unresponsive)&lt;br&gt;
Forza Horizon 6 — GPU hangs tied to the VKD3D-Proton descriptor heap feature&lt;br&gt;
Elden Ring and Elden Ring Nightreign&lt;br&gt;
Plus fixes touching Assassin's Creed Origins, Total War: Warhammer III, Crimson Desert, X-Plane, and a long tail of smaller titles (007 First Light, Grounded 2, Star Rupture, Windrose, and others)&lt;/p&gt;

&lt;h2&gt;
  
  
  Outside of gaming fixes, two changes matter for general desktop stability:
&lt;/h2&gt;

&lt;p&gt;A fix for an X Server crash under GLX indirect rendering&lt;br&gt;
A fix for black screens after modesets in X11 apps using the Present extension&lt;/p&gt;

&lt;p&gt;For anyone running a Turing-generation GPU or newer, installation has also gotten noticeably less painful — modeset configuration is now handled automatically, closer to a plug-and-play experience than the manual dance NVIDIA-on-Linux has traditionally required.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should you actually install it?
&lt;/h2&gt;

&lt;p&gt;If you're not chasing bleeding-edge features and just want the most predictable driver, 595.91.07 (the parallel stable release) is the safer pick. But if any of the games above have been giving you trouble, or you've been holding off on an update because previous 610-branch builds felt rough, this specific point release is the one that closes out most of those complaints — 610.57.04 was the first release in the R610 feature branch back in late May, and this is the point release that actually fixes what broke on launch.&lt;/p&gt;

&lt;p&gt;The installer is available directly from NVIDIA (~442 MB), with DKMS, RPM, and deb packages supported out of the box. If you're new to NVIDIA-on-Linux, installing through the .run file is still not the recommended path for beginners — wait for your distro's packaged version if you can, since it's much harder to break.&lt;/p&gt;

&lt;h2&gt;
  
  
  Full verdict and version comparison
&lt;/h2&gt;

&lt;p&gt;I run this driver daily on my own homelab/gaming setup, so I put together a deeper breakdown — full version comparison table, install notes, and my actual verdict after running it for real — over on El Rack:&lt;/p&gt;

&lt;p&gt;👉 &lt;a href="https://elrack.es/gaming/nvidia-driver-610-57-04-linux/" rel="noopener noreferrer"&gt;Driver NVIDIA 610.57.04 para Linux&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;(Spanish-language site — the specs table, version numbers, and rating are readable regardless, and translation tools handle the rest cleanly if you want the full writeup.)&lt;/p&gt;

&lt;p&gt;I write about homelab, self-hosting, and Linux gaming at El Rack — real testing from inside my own homelab, not just a spec sheet rewrite.&lt;/p&gt;

</description>
      <category>nvidia</category>
      <category>linux</category>
      <category>gaming</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Claude Code Now Runs Subagents in the Background by Default — What Actually Changed</title>
      <dc:creator>Alvarito1983</dc:creator>
      <pubDate>Fri, 07 Aug 2026 07:15:58 +0000</pubDate>
      <link>https://dev.to/alvarito1983/claude-code-now-runs-subagents-in-the-background-by-default-what-actually-changed-54kb</link>
      <guid>https://dev.to/alvarito1983/claude-code-now-runs-subagents-in-the-background-by-default-what-actually-changed-54kb</guid>
      <description>&lt;p&gt;Original article in Spanish on El Rack: Claude Code ya no espera a que le mires: subagentes en segundo plano por defecto&lt;/p&gt;

&lt;p&gt;I run Claude Code daily as my main dev agent, both locally and over SSH against a VPS running a handful of production sites. A change that landed over the last few weeks quietly rewired how I work with it: subagents now run in the background by default.&lt;/p&gt;

&lt;p&gt;Before, when Code delegated a task to a subagent, the conversation blocked until that subagent finished. Since July 1, it keeps working on other things while the subagent runs, and notifies you when it's done. For anyone running long sessions with &lt;code&gt;--dangerously-skip-permissions&lt;/code&gt; and checking back hours later, that's the difference between a tool that waits and one that actually works unattended.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes day to day
&lt;/h2&gt;

&lt;p&gt;My workflow has always been the same: approve the initial plan, then let Code proceed without confirmation on every step, stopping only on real errors. With background subagents, that autonomy compounds. I can now ask it to audit a site, generate new content, and review config files at the same time — each running as an independent subagent instead of queuing one after another.&lt;/p&gt;

&lt;p&gt;Two limits matter here: a default cap of 20 concurrent subagents, and nesting now allowed up to 3 levels deep (up from 1). Both exist for a reason — it's easy for a large task to branch further than you expect if nothing constrains it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sonnet 5 and Opus 5: 1M context is no longer the exception
&lt;/h2&gt;

&lt;p&gt;Claude Sonnet 5 became the default model in late June, with a native 1M-token context window. Claude Opus 5 replaced it as the default Opus model on July 24, also with 1M context. In long sessions — the kind you get during a full editorial rebuild of a site — I used to watch &lt;code&gt;/context&lt;/code&gt; and run &lt;code&gt;/compact&lt;/code&gt; mid-task to avoid losing thread. Now I run noticeably longer sessions without touching either. It's not magic: auto-compaction still kicks in, and cost on very long Opus sessions is worth watching, but the practical ceiling moved.&lt;/p&gt;

&lt;h2&gt;
  
  
  The security fix that actually matters
&lt;/h2&gt;

&lt;p&gt;As someone who manages production infrastructure, this is the line item I care about most: in August, a fix landed for worktree-isolated sessions (and their subagents) that could run destructive git commands against the main checkout. When you're launching subagents with broad permissions against something that matters, broken isolation like that is exactly the failure mode you don't want. That it got closed quickly is a good sign — but also a reminder that "autonomous" and "unsupervised" aren't the same thing. I still review the final report of every session rather than assuming no visible errors means everything went well.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I still draw lines
&lt;/h2&gt;

&lt;p&gt;With this much autonomy, the temptation is to hand off increasingly ambitious tasks and walk away. I don't do that on anything touching client-facing production servers — those sessions stay interactive, with explicit confirmation on sensitive steps. Where I do give it full rein is on my own infrastructure, where the cost of a mistake is low and recoverable. That distinction, more than any config flag, is what actually determines how much autonomy a session gets.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;Background-by-default subagents are the change that's altered my day-to-day with Claude Code the most in months — not because it's flashy, but because it fits how I was already using it: approve the plan, let it run. The 1M context on Sonnet 5 and Opus 5 supports that well. The worktree isolation fix is the reminder that more autonomy means reviewing final output more carefully, not less.&lt;/p&gt;




&lt;p&gt;Original (Spanish, full version with pros/cons and more context): &lt;a href="https://dev.tourl"&gt;&lt;/a&gt;&lt;a href="https://elrack.es/herramientas-ia/claude-code-subagentes-segundo-plano-opus-5-2026/" rel="noopener noreferrer"&gt;https://elrack.es/herramientas-ia/claude-code-subagentes-segundo-plano-opus-5-2026/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>claude</category>
      <category>llm</category>
    </item>
    <item>
      <title>Kimi K3 is the largest open-weight model ever released — and you probably still can't run it</title>
      <dc:creator>Alvarito1983</dc:creator>
      <pubDate>Thu, 06 Aug 2026 06:45:20 +0000</pubDate>
      <link>https://dev.to/alvarito1983/kimi-k3-is-the-largest-open-weight-model-ever-released-and-you-probably-still-cant-run-it-1nn3</link>
      <guid>https://dev.to/alvarito1983/kimi-k3-is-the-largest-open-weight-model-ever-released-and-you-probably-still-cant-run-it-1nn3</guid>
      <description>&lt;p&gt;Originally published in Spanish on El Rack. Browser translation handles the rest of the site fine if you're into homelab/self-hosting content.&lt;/p&gt;

&lt;p&gt;Moonshot AI released Kimi K3 on July 17, 2026, and made the weights publicly downloadable on July 27. At 2.8 trillion parameters, it's the largest open-weight model ever published — and according to multiple benchmarks, it rivals Claude Opus and GPT on coding, reasoning, and general knowledge work, at a fraction of the training cost.&lt;/p&gt;

&lt;p&gt;The New York Times ran an in-depth piece on it a few days after release, which tells you this isn't just another model drop.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "open weights" actually gets you here
&lt;/h2&gt;

&lt;p&gt;Publicly downloadable weights mean any company or researcher can run this locally and modify it without depending on a third-party API. If you already run Ollama or LM Studio in your homelab, that's the tempting part: a frontier-level model, no monthly quota, running on your own hardware.&lt;/p&gt;

&lt;h2&gt;
  
  
  The practical reality is different.
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;"2.8 trillion parameters isn't a number that runs on homelab hardware — it needs an enterprise-grade GPU cluster. The weight release is real, but "downloadable" and "runnable" are very different things at this scale."&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  The bigger debate this reopened
&lt;/h2&gt;

&lt;p&gt;What makes Kimi K3 interesting isn't just the benchmark numbers — it's what it represents in the ongoing dispute over AI's geopolitics. The same fracture that opened up around DeepSeek-R1 in January 2025 is back: some argue US labs need to close up more in response to Chinese competition, others see openness as the only real way to stay relevant against an ecosystem that ships open weights at a pace closed labs can't match on transparency. There's also a real technical concern underneath: the possibility that outside actors use massive querying of closed American models to distill their outputs and train competing open models.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this actually matters for a homelab
&lt;/h2&gt;

&lt;p&gt;Even though K3 itself is unrunnable on consumer hardware, its release pushes down what smaller, actually-runnable models (7B-70B, the ones that fit on a consumer GPU) can eventually achieve — research and techniques from frontier releases like this tend to filter down, via distillation, into much more manageable versions. Ollama's ecosystem usually adds support for distilled variants of these releases within days or weeks.&lt;/p&gt;




&lt;h2&gt;
  
  
  Verdict
&lt;/h2&gt;

&lt;p&gt;Kimi K3 isn't something you're installing in your homelab this week, but it's a meaningful signal of where open AI is heading — and probably the source of much smaller, distilled versions that will show up in Ollama soon. Worth tracking not for what you can run today, but for what it anticipates for the near future.&lt;/p&gt;

&lt;p&gt;Full article (in Spanish): &lt;a href="https://dev.tourl"&gt;&lt;/a&gt; &lt;a href="https://elrack.es/herramientas-ia/kimi-k3/" rel="noopener noreferrer"&gt;https://elrack.es/herramientas-ia/kimi-k3/&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>opensource</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>The MCP servers actually worth installing in Claude Code (2026) Published</title>
      <dc:creator>Alvarito1983</dc:creator>
      <pubDate>Wed, 05 Aug 2026 07:15:44 +0000</pubDate>
      <link>https://dev.to/alvarito1983/the-mcp-servers-actually-worth-installing-in-claude-code-2026published-43lb</link>
      <guid>https://dev.to/alvarito1983/the-mcp-servers-actually-worth-installing-in-claude-code-2026published-43lb</guid>
      <description>&lt;p&gt;Originally published in Spanish on El Rack, a Spanish tech review site I write for. Browser translation handles the rest of the site well if you want to dig deeper into the sysadmin/homelab content.&lt;/p&gt;

&lt;p&gt;Claude Code is my main tool every day for managing a handful of self-hosted projects. So when "just install these 10 MCP servers" guides kept showing up in every forum thread, I wanted to actually separate what's worth the setup time from what's just noise sitting in your context window.&lt;/p&gt;

&lt;p&gt;Here's what I found after digging into the current state of the ecosystem.&lt;/p&gt;

&lt;p&gt;The cost nobody mentions before you install ten of these&lt;/p&gt;

&lt;p&gt;Before the list: the one number that matters most. Every connected MCP server injects somewhere between 2,000 and 5,000 tokens of tool schemas at session start. Three servers connected at once is already 6,000 to 15,000 tokens spent before you type your first prompt.&lt;/p&gt;

&lt;p&gt;That's not free, even though it feels like it. If Claude Code feels slower to start or seems to "forget" context earlier than you'd expect, check how many MCP servers you're leaving connected by default.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Correction (added after publishing): Claude Code now enables Tool Search by default, which defers MCP schema loading until a tool is actually used — a connected-but-idle server now costs closer to a name-and-description stub than a full schema dump. The 2,000–5,000 token figure above was accurate before this became the default, and still applies if you disable Tool Search or if a given server doesn't support deferred loading yet. Thanks to a reader for catching this in the comments.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;MCP is convenient. It is not free. Every server you connect is tokens that are no longer available for your own code.&lt;/p&gt;

&lt;p&gt;The three actually worth installing first&lt;/p&gt;

&lt;p&gt;Independent guides keep converging on the same top three, and after using them daily, I agree with the consensus:&lt;/p&gt;

&lt;p&gt;Context7 — live library documentation, so Claude doesn't hallucinate APIs that changed three versions ago. This is the one that's saved me the most from confidently-wrong code suggestions based on outdated library docs.&lt;/p&gt;

&lt;p&gt;Official GitHub MCP — repo operations, pull requests, issues and CI, without leaving the chat. One caveat on the token cost above: for simple lookups, the gh CLI is still much cheaper than going through MCP.&lt;/p&gt;

&lt;p&gt;Playwright MCP — lets Claude drive a real browser to confirm a frontend change actually renders correctly, instead of just trusting the code diff. If you do any frontend work, this is the difference between "the code compiles" and "the code actually works."&lt;/p&gt;

&lt;p&gt;Two that Anthropic still maintains but aren't essential anymore&lt;/p&gt;

&lt;p&gt;The original reference servers for structured reasoning and file operations are still around, but the general consensus is honest about it: native model reasoning improved enough through 2026 that most people skip the reasoning server entirely, and Claude Code already ships built-in file tools that make the equivalent MCP mostly redundant.&lt;/p&gt;

&lt;p&gt;The governance shift that matters long-term&lt;/p&gt;

&lt;p&gt;In December 2025, MCP governance moved from Anthropic alone to the Agentic AI Foundation under the Linux Foundation, with Anthropic, OpenAI, Google, Microsoft, AWS, Cloudflare and Bloomberg on the board. In practice, this means vendor-maintained servers (Slack's own MCP, GitHub's own MCP) tend to be more reliable than generic third-party integrations — each company simply knows its own service better than any middleman.&lt;/p&gt;

&lt;p&gt;Best MCP for a specific task&lt;/p&gt;

&lt;p&gt;If you don't want to read the full ranking, here's the direct answer for whatever you're trying to solve:&lt;/p&gt;

&lt;p&gt;Working with your database → The PostgreSQL MCP if you run Postgres on any provider, or the official Supabase MCP if your backend lives there specifically (you can scope it to a single project via a URL parameter — good security practice).&lt;/p&gt;

&lt;p&gt;Repo management → The official GitHub MCP, already covered above.&lt;/p&gt;

&lt;p&gt;Verifying frontend actually works → Playwright, also covered above.&lt;/p&gt;

&lt;p&gt;Project docs and knowledge base → The Notion MCP, if your team documents there — keeps pages synced with code changes automatically.&lt;/p&gt;

&lt;p&gt;Task management → The Linear MCP, the most widely adopted for this.&lt;/p&gt;

&lt;p&gt;Production error monitoring → The Sentry MCP — brings real production error context straight into the chat instead of copy-pasting logs by hand.&lt;/p&gt;

&lt;p&gt;Design → The Figma MCP — pulls measurements, components and assets directly from published designs.&lt;/p&gt;

&lt;p&gt;Notifying a team → The Slack MCP — post deployment updates or task results without leaving the session.&lt;/p&gt;

&lt;p&gt;The question isn't "which MCP servers exist" — it's "which task do I repeat constantly that could be solved without switching windows." That question gets you a better answer than any generic ranking.&lt;/p&gt;

&lt;p&gt;When to use MCP and when not to&lt;/p&gt;

&lt;p&gt;The rule that sums it up best: use a CLI when latency and tokens matter and you can make a direct shell call; use an MCP server when you want the same tool available across Claude Code, Cursor and other clients with one config; use a Skill when the workflow is a set of instructions, not an external tool to call.&lt;/p&gt;

&lt;p&gt;Full article (in Spanish) with more detail: elrack.es/herramientas-ia/mejores-mcp-claude-code-2026&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claudecode</category>
      <category>mcp</category>
      <category>productivity</category>
    </item>
    <item>
      <title>I asked Claude Code to refactor a feature it built last week. It hated its own code.</title>
      <dc:creator>Alvarito1983</dc:creator>
      <pubDate>Mon, 27 Apr 2026 08:50:42 +0000</pubDate>
      <link>https://dev.to/alvarito1983/i-asked-claude-code-to-refactor-a-feature-it-built-last-week-it-hated-its-own-code-ep8</link>
      <guid>https://dev.to/alvarito1983/i-asked-claude-code-to-refactor-a-feature-it-built-last-week-it-hated-its-own-code-ep8</guid>
      <description>&lt;p&gt;Last week I opened a feature I'd built six days earlier and asked Claude Code to extend it.&lt;/p&gt;

&lt;p&gt;I hadn't touched the file since. The code was working. Tests passing. Nothing was on fire. I just wanted to add a small capability — fifteen minutes of work, maybe.&lt;/p&gt;

&lt;p&gt;What I got back wasn't an extension. It was a quiet roast.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Before extending this, I'd recommend refactoring. The current implementation has a few issues: the error handling is inconsistent, the function is doing more than its name suggests, and there's a duplicated validation block that should be a helper. Want me to fix those first?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Reader, that code was written by Claude Code. Same project. Same model tier. Six days earlier. By the same agent that was now telling me, with polite confidence, that it was a mess.&lt;/p&gt;

&lt;p&gt;I sat with that for a minute. Then I started thinking about what it meant. Because the more I thought, the more I realized that this — this exact moment — is the thing nobody is writing about with AI coding agents.&lt;/p&gt;




&lt;h2&gt;
  
  
  The "memory problem" everyone complains about
&lt;/h2&gt;

&lt;p&gt;If you've used Claude Code, Codex, Cursor in agent mode, or any of the rest, you've hit it: agents don't have continuity between sessions. Every new conversation starts cold. They don't remember what they decided last Tuesday. They don't remember why a particular abstraction exists. They don't remember that you already debated naming this function and settled on something specific.&lt;/p&gt;

&lt;p&gt;Most of the discourse treats this as a problem to solve. We bolt on memory systems. We pile up CLAUDE.md files. We build context-loading scripts. The implicit assumption is that the agent &lt;em&gt;should&lt;/em&gt; remember, and the work is to compensate for the fact that it doesn't.&lt;/p&gt;

&lt;p&gt;I'm starting to think that framing is exactly backwards.&lt;/p&gt;




&lt;h2&gt;
  
  
  What actually happened in that session
&lt;/h2&gt;

&lt;p&gt;When Claude Code criticized its own code from six days earlier, it wasn't being inconsistent. It wasn't malfunctioning. It was being &lt;em&gt;honest&lt;/em&gt; — in a way no human collaborator can.&lt;/p&gt;

&lt;p&gt;Think about what just happened:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A reviewer with no ego attachment to the code looked at it fresh.&lt;/li&gt;
&lt;li&gt;It had no memory of why we wrote it that way, what we'd ruled out, what compromises we made under time pressure, or what I muttered at 11pm when the test finally passed.&lt;/li&gt;
&lt;li&gt;It just looked at the code on its merits and said: &lt;em&gt;this could be better.&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's not a bug. That's the cleanest code review you'll ever get.&lt;/p&gt;

&lt;p&gt;A human reviewing their own code from a week ago has every cognitive bias working against them. They remember the constraints. They remember the deadline. They remember that they "knew" the helper was duplicated but didn't want to break the function signature. They protect their past decisions because their past decisions are &lt;em&gt;theirs&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Claude Code, with no continuity, has none of that. Every session is a new pair of eyes that hasn't been bought off by yesterday's reasoning.&lt;/p&gt;




&lt;h2&gt;
  
  
  The reframe
&lt;/h2&gt;

&lt;p&gt;Once I saw this, I started doing it on purpose.&lt;/p&gt;

&lt;p&gt;When I finish a feature now, I don't immediately ship it. I commit, sleep on it, and the next morning I open a new Claude Code session and ask it to review the code as if it had never seen it. Because for that session, it hasn't.&lt;/p&gt;

&lt;p&gt;The reviews are brutal. They are also, almost always, right.&lt;/p&gt;

&lt;p&gt;It catches:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Functions that are doing two things and lying about it in their name&lt;/li&gt;
&lt;li&gt;Error handling that's defensive in some places and absent in others&lt;/li&gt;
&lt;li&gt;Helpers I should have extracted but didn't&lt;/li&gt;
&lt;li&gt;Naming that made sense in context but doesn't survive cold reading&lt;/li&gt;
&lt;li&gt;Tests that pass but verify the wrong thing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The same agent that wrote the code can't see these issues &lt;em&gt;while it's writing&lt;/em&gt;, because it has the same tunnel vision a human does in flow. But strip its memory of the session, hand it the file, and it becomes the reviewer it could never be in real-time.&lt;/p&gt;

&lt;p&gt;The "memory problem" is actually a feature gate to a different mode of useful behavior. We've been trying to remove it instead of using it.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why this changes how I think about CLAUDE.md
&lt;/h2&gt;

&lt;p&gt;I've written about CLAUDE.md before, and I want to refine what I said there, because this experience clarified something for me.&lt;/p&gt;

&lt;p&gt;CLAUDE.md is not the agent's memory. That framing is wrong.&lt;/p&gt;

&lt;p&gt;CLAUDE.md is the project's memory of decisions that shouldn't be relitigated every session. Things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"We use Tailwind v3, not v4. Don't migrate."&lt;/li&gt;
&lt;li&gt;"The auth layer is intentionally simple. Don't add OAuth without asking."&lt;/li&gt;
&lt;li&gt;"Tests live next to source files, not in /tests."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are &lt;em&gt;decisions&lt;/em&gt;, not &lt;em&gt;style&lt;/em&gt;. CLAUDE.md exists so the agent doesn't waste a session arguing with itself about settled choices.&lt;/p&gt;

&lt;p&gt;But CLAUDE.md should &lt;em&gt;not&lt;/em&gt; contain things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"This function does X" (the code says that)&lt;/li&gt;
&lt;li&gt;"Last week we refactored Y" (that's history, not a decision)&lt;/li&gt;
&lt;li&gt;"The pattern we use is Z" (let the code be the source of truth)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The mistake is using CLAUDE.md as a memory dump. When you do that, you're trying to give the agent continuity — and you lose the fresh-eyes review benefit. You've turned your honest reviewer into a yes-man who already agrees with everything you've decided.&lt;/p&gt;

&lt;p&gt;The right CLAUDE.md is short. It pins down decisions. It leaves &lt;em&gt;judgment&lt;/em&gt; alone, so a new session can still tell you when your code is bad.&lt;/p&gt;




&lt;h2&gt;
  
  
  The mental model shift
&lt;/h2&gt;

&lt;p&gt;Stop thinking of an AI coding agent as one consistent collaborator across time.&lt;/p&gt;

&lt;p&gt;Start thinking of it as a high-quality reviewer you can summon, fresh, as many times as you want — provided you preserve the decisions, not the reasoning.&lt;/p&gt;

&lt;p&gt;This changes the workflow:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Write code with the agent in a session.&lt;/li&gt;
&lt;li&gt;Commit when it works.&lt;/li&gt;
&lt;li&gt;Open a new session the next day. Ask it to review.&lt;/li&gt;
&lt;li&gt;Take the criticism seriously. It's coming from someone who isn't defending yesterday's choices.&lt;/li&gt;
&lt;li&gt;If the criticism is wrong, that's a signal that a &lt;em&gt;decision&lt;/em&gt; is missing from CLAUDE.md.&lt;/li&gt;
&lt;li&gt;If the criticism is right, fix it. The fact that the same agent wrote the code is irrelevant.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I've started shipping noticeably cleaner code since I adopted this. Not because Claude Code got smarter. Because I stopped trying to glue together its sessions and started using the gaps between them.&lt;/p&gt;




&lt;h2&gt;
  
  
  The closing thought
&lt;/h2&gt;

&lt;p&gt;There's a lot of energy right now going into making AI agents more continuous, more memory-rich, more "aware" across sessions. Long-running agents. Persistent context. Memory layers.&lt;/p&gt;

&lt;p&gt;I'm not against any of that. But I think the rush is hiding something valuable. The discontinuity between sessions is the thing that lets the agent be a real reviewer instead of a sympathetic teammate. Once you give it perfect memory, it stops being able to look at your code from outside, and it starts protecting your past decisions the way a human would.&lt;/p&gt;

&lt;p&gt;Maybe the goal isn't to give agents continuity. Maybe it's to give them just enough of it — your settled decisions — and protect the rest. Let them forget how the code got that way. Let them tell you, fresh, that it's bad.&lt;/p&gt;

&lt;p&gt;Mine did. It was right.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Has Claude Code ever roasted its own past output in your projects? I want to hear the worst one — drop it in the comments.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>tooling</category>
      <category>webdev</category>
    </item>
    <item>
      <title>The AI Coding Tools Panorama in 2026: From Claude Code to the Free Alternatives That Actually Replace It</title>
      <dc:creator>Alvarito1983</dc:creator>
      <pubDate>Sat, 25 Apr 2026 16:18:39 +0000</pubDate>
      <link>https://dev.to/alvarito1983/the-ai-coding-tools-panorama-in-2026-from-claude-code-to-the-free-alternatives-that-actually-1p0a</link>
      <guid>https://dev.to/alvarito1983/the-ai-coding-tools-panorama-in-2026-from-claude-code-to-the-free-alternatives-that-actually-1p0a</guid>
      <description>&lt;p&gt;The AI coding tool space in 2026 looks nothing like it did 18 months ago. Autocomplete is a solved problem. The interesting question now is: which agent do you trust to read your codebase, plan a refactor, run your tests, and not torch your API budget while doing it.&lt;/p&gt;

&lt;p&gt;I've been using these tools daily for the last year on a real project — a Docker management suite I'm building solo, where the cost of a bad refactor is hours of cleanup. That context matters, because most "best AI coding tools" lists rank by benchmark scores. Benchmarks measure isolated tasks. They don't measure what happens at hour three of a complex migration when the agent forgets which files it already touched.&lt;/p&gt;

&lt;p&gt;This is the honest panorama from someone shipping production code. From the ones everyone knows to the ones that quietly outperform them. Then two opinionated rankings: the five paid tools that justify their price, and the five free ones that make you question whether you need to pay at all.&lt;/p&gt;

&lt;p&gt;Let's get into it.&lt;/p&gt;




&lt;h2&gt;
  
  
  The full panorama
&lt;/h2&gt;

&lt;p&gt;The space has split into clear layers. Most developers running serious AI workflows use 2–3 tools, not one — and once you understand the layers, the picking gets easier.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 1: Inline assistants (autocomplete + chat in your editor)
&lt;/h3&gt;

&lt;p&gt;These are what most people still think "AI coding" means. They suggest the next line, complete a function, answer a quick question. Low cognitive overhead, high frequency, low ambition.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub Copilot&lt;/strong&gt; is the default. 76% of developers have heard of it, 29% use it at work, and at $10/month for Pro it's the cheapest entry point that doesn't feel like a compromise. It ships in every editor that matters and just works. The growth has stalled — adoption has plateaued — but it's still the safest choice for someone who wants AI in their editor and doesn't want to think about it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;JetBrains AI Assistant + Junie&lt;/strong&gt; is the equivalent for the JetBrains crowd. 11% adoption combined. If you live in IntelliJ, PyCharm, or WebStorm, it's the natural fit because it understands JetBrains' code intelligence in ways Copilot doesn't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tabnine&lt;/strong&gt; still exists and still gets recommended for teams that need on-prem or air-gapped deployment. Outside that niche, it's been overtaken.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Continue.dev&lt;/strong&gt; is the open-source Copilot-shaped option. Lives inside VS Code or JetBrains as an extension, brings your own API key, 31K GitHub stars. Less polished than Copilot, but you're not on a per-seat license and you choose your model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 2: AI-native IDEs (the editor itself is the product)
&lt;/h3&gt;

&lt;p&gt;These are full editor replacements where the AI is woven into every keystroke, not bolted on as an extension.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cursor&lt;/strong&gt; is the reference. A VS Code fork that took "AI in the editor" further than anyone else. Composer mode handles multi-file edits with visual diffs, autocomplete is supernaturally fast (Supermaven under the hood), and the agent can take on actual tasks. It's the most productive IDE-based AI experience right now, with one giant asterisk: pricing trust took a hit earlier in 2026 when Anthropic's pricing changes cascaded through Cursor's billing model and a lot of people got surprise bills. The product is still excellent. The pricing is the part you have to watch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Windsurf&lt;/strong&gt; is the value alternative. Same category as Cursor, $15/month base, free tier with full IDE features and the Cascade agent. It's been climbing the rankings precisely because it's what Cursor was before the pricing got messy.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Antigravity&lt;/strong&gt; is Google's entry, launched November 2025. Free during preview. Supports Claude, Gemini, GPT-OSS — the most diverse model lineup of any free tool right now. 6% adoption already, which is fast for a tool this new.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 3: Terminal agents (the new center of gravity)
&lt;/h3&gt;

&lt;p&gt;This is where serious work is happening in 2026. You point an agent at your repo from the terminal, and it reads, edits, runs tests, iterates. The terminal-first approach composes with everything: git, your shell, your CI, your existing scripts.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude Code&lt;/strong&gt; is the one I use daily and the one most experienced developers have settled on. Anthropic's official terminal agent, runs Opus and Sonnet, scores 80.8% on SWE-bench Verified — meaning it actually solves real GitHub issues, not toy problems. It's the best at multi-file reasoning and the best at not losing context on hour three of a complex task. Costs $20/month for Pro, but heavy use can run $100–200/month on the Max plan or via API. That's the elephant in the room.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OpenAI Codex CLI&lt;/strong&gt; re-entered the conversation in early 2026 with parallel sandboxed execution and automatic PR creation. 3% adoption (data from before its desktop app launched), but climbing. Strong choice if you're already in the OpenAI ecosystem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gemini CLI&lt;/strong&gt; is the underrated one. Google's terminal agent. 1,000 free requests/day with Gemini 2.5 Pro and a 1M context window. Less consistent than Claude on complex refactors, but for the price (zero) and the context size (massive), it's the best free terminal option right now. Don't skip it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Aider&lt;/strong&gt; is the elder statesman. 41K GitHub stars, terminal-first, and its defining feature is that every AI edit is a git commit. You get a complete audit trail of what the AI changed and why. Bring your own API key. If you live in git, this is the one to try first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OpenCode&lt;/strong&gt; is the most popular open-source terminal agent. 95K stars. Provider-agnostic — supports 75+ models. Free models included, plus you can plug in any API key. Cleaner TUI than Aider, less opinionated about git.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cline&lt;/strong&gt; (58K stars) and its forks &lt;strong&gt;Roo Code&lt;/strong&gt; (22K) and &lt;strong&gt;Kilo Code&lt;/strong&gt; (16K) live as VS Code extensions but operate as full agents — they edit files, run commands, and ask for approval at each step. BYOK with zero markup. If you want Cursor-style agentic work but inside vanilla VS Code without the subscription, this is the path.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 4: Cloud agents (autonomous, parallel, expensive)
&lt;/h3&gt;

&lt;p&gt;These run in the cloud, often handling multiple tasks in parallel sandboxes, opening PRs against your repo while you do something else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Devin&lt;/strong&gt; is the original autonomous agent. 67% PR merge rate on well-defined tasks. Treats coding tasks like Jira tickets it picks up and ships. $20/month base plus unpredictable per-task costs. Useful as an accelerator for narrow, well-scoped work; treat anything it ships as a draft.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Codex (cloud version)&lt;/strong&gt; is OpenAI's hosted agent that runs in sandboxed environments and pushes PRs to GitHub. Strong if your stack is OpenAI-aligned.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;OpenHands&lt;/strong&gt; (formerly OpenDevin) is the open-source version of this category. 68K stars, MIT-licensed, BYOK. If you want a cloud-style autonomous agent without the SaaS dependency, this is it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 5: Specialized tools (the rest)
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Replit&lt;/strong&gt; still owns the "build an app from a prompt in your browser" niche, especially for prototyping and education. Costs accumulate fast at scale.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bolt.new&lt;/strong&gt; and &lt;strong&gt;Lovable.dev&lt;/strong&gt; are similar — browser-based, AI-first, great for MVPs and demos. Lovable couples to Supabase tightly, which is convenient until you outgrow it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;v0&lt;/strong&gt; (Vercel) generates React components from prompts or Figma designs. Useful for design-to-code, less so for general work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tabby&lt;/strong&gt; is self-hosted autocomplete you run on your own GPU. The privacy-first option for teams that can't send code to anyone.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Snyk Code&lt;/strong&gt; and &lt;strong&gt;Qodo&lt;/strong&gt; sit alongside the rest of the stack — security scanning and AI code review on PRs. Not coding tools strictly, but they're part of the modern AI workflow.&lt;/p&gt;




&lt;h2&gt;
  
  
  Top 5 paid tools that actually justify their price
&lt;/h2&gt;

&lt;p&gt;I'm ranking these on a single criterion: would I be measurably less productive without them. Not "is the demo impressive." Not "did they raise a Series B." Would my output drop if I uninstalled it tomorrow.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Claude Code — $20–200/month
&lt;/h3&gt;

&lt;p&gt;Best for: complex multi-file work, anything where the model losing context costs you an hour of cleanup.&lt;/p&gt;

&lt;p&gt;Claude Code on Opus is the only tool I trust with a refactor that touches more than three files. It plans, executes, runs my tests, and recovers from its own mistakes. The 200K context window plus extended thinking actually changes how I architect things — I can describe a problem at a higher level than I used to.&lt;/p&gt;

&lt;p&gt;The price is real. Heavy use can run $100–200/month on Max. The honest framing is: if you're a working developer billing for your time, $200/month for the tool that saves you a few hours a week is the cheapest line item on your invoice. If you're a hobbyist, this isn't the right tier.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Cursor — from $20/month
&lt;/h3&gt;

&lt;p&gt;Best for: developers who think visually and want diffs they can scan instead of accept.&lt;/p&gt;

&lt;p&gt;If you do most of your work in an editor and the terminal-first model doesn't fit your brain, Cursor is the most productive AI IDE in existence. Composer mode for multi-file changes, instant autocomplete, model orchestration. The pricing trust issue I mentioned is real — monitor your credits, especially if you flip into agent mode for big tasks. With that caveat, it earns its place here.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. GitHub Copilot — $10/month
&lt;/h3&gt;

&lt;p&gt;Best for: the developer who wants AI in their editor and wants to never think about it again.&lt;/p&gt;

&lt;p&gt;Copilot earned its place not by being the best, but by being the floor. $10/month, works everywhere, never surprises you with a bill, integrates with the GitHub ecosystem (PR summaries, issue context, repo activity). It's the silent default. Stop overthinking.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Windsurf — from $15/month
&lt;/h3&gt;

&lt;p&gt;Best for: people who want what Cursor was before the pricing drama.&lt;/p&gt;

&lt;p&gt;Windsurf does what Cursor does, charges less, and the free tier is genuinely usable for daily work. Cascade agent, plan mode, parallel multi-agent sessions with git worktrees. It's the IDE I'd recommend to someone starting from scratch in 2026 if cost predictability matters.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Devin — $20/month + variable
&lt;/h3&gt;

&lt;p&gt;Best for: well-scoped, repetitive tasks you can describe in a paragraph.&lt;/p&gt;

&lt;p&gt;Devin shipping a 67% PR merge rate on defined tasks is the data point that matters. It's not autonomy in the science-fiction sense — it's a junior engineer who works on tickets in parallel sandboxes and submits drafts. For the right kind of work (boring but real: dependency upgrades, test coverage, scoped refactors) it earns its keep. For ambiguous work, you'll spend more time correcting it than you save.&lt;/p&gt;




&lt;h2&gt;
  
  
  Top 5 free tools that genuinely replace a paid one
&lt;/h2&gt;

&lt;p&gt;The bar here: would I recommend this to someone who can't or won't pay, knowing they'll get within 80% of the paid experience.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Gemini CLI — free (1,000 requests/day)
&lt;/h3&gt;

&lt;p&gt;Best free terminal agent in 2026, full stop. Gemini 2.5 Pro, 1M context window, runs in your terminal, BYOK with a generous free tier from Google. It's not as consistent as Claude Code on the hardest tasks, but for 90% of work, the gap doesn't matter. And the 1M context window means it can ingest a small codebase in a single prompt — something even paid Claude Code can't always match.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Aider — free (BYOK)
&lt;/h3&gt;

&lt;p&gt;Best for terminal-native, git-centric work. Pair it with the DeepSeek API and you're paying $5–15/month total for AI coding that competes with $200/month tools. Every edit is a git commit. Reviewable, revertible, auditable. If you've ever lost track of what an agent changed across a session, Aider's commit history is the answer.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Cline — free (BYOK, runs in VS Code)
&lt;/h3&gt;

&lt;p&gt;Best free agent inside an IDE. 5M+ installs, Apache-licensed, every action requires human approval. Plug in your Claude or OpenAI key and you have Cursor's agent capability inside vanilla VS Code without the subscription. The forks (Roo Code, Kilo Code) add structured modes and broader model support — pick whichever feels right; they're all good.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. GitHub Copilot Free — free (2,000 completions + 50 chats/month)
&lt;/h3&gt;

&lt;p&gt;The free tier of Copilot is real and surprisingly generous for casual or learning use. If you're not coding 8 hours a day, the free quota covers it. The path most people should take: start here, find your edges, then decide whether to pay or move to BYOK.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. OpenCode — free (BYOK, free models included)
&lt;/h3&gt;

&lt;p&gt;The open-source terminal agent that's quietly become the most popular on GitHub (95K stars). Ships with free models you can use immediately, supports 75+ providers when you bring keys, polished TUI. If you want to try terminal agents without making any decisions about API keys or pricing on day one, OpenCode is the lowest-friction starting point.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I actually run
&lt;/h2&gt;

&lt;p&gt;For full transparency, since I think it matters more than abstract rankings: I use Claude Code on Max for the heavy lifting, Cline in VS Code for in-editor work where I want approval gates, and Gemini CLI for anything where I want to throw a huge codebase at a model in one prompt. Three tools, three roles, no overlap.&lt;/p&gt;

&lt;p&gt;That's the meta-point. In 2026, the question is no longer "which is the best AI coding tool." It's "which combination handles each layer of my workflow without breaking under pressure." Get that right and the cost question answers itself — because the right setup pays for itself in saved hours, and the wrong setup torches credits without shipping anything.&lt;/p&gt;

&lt;p&gt;If your current setup is one tool doing everything, you're probably either overpaying for capability you don't use or underpowered for the work you actually do. Pick the layer that hurts most and start there.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;What does your stack look like? Curious especially about people running fully BYOK setups — what's the monthly bill landing at versus the equivalent SaaS subscriptions?&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>tooling</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
