<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Juber Shaikh</title>
    <description>The latest articles on DEV Community by Juber Shaikh (@zubair_shaikh_a22cea567b0).</description>
    <link>https://dev.to/zubair_shaikh_a22cea567b0</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1718777%2F9cd764ae-3c11-4035-b1a3-f2e031e30415.jpg</url>
      <title>DEV Community: Juber Shaikh</title>
      <link>https://dev.to/zubair_shaikh_a22cea567b0</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/zubair_shaikh_a22cea567b0"/>
    <language>en</language>
    <item>
      <title>I stopped trusting my AI agents' "done" — here's the boring gate that fixed it</title>
      <dc:creator>Juber Shaikh</dc:creator>
      <pubDate>Tue, 29 Sep 2026 04:53:43 +0000</pubDate>
      <link>https://dev.to/zubair_shaikh_a22cea567b0/i-stopped-trusting-my-ai-agents-done-heres-the-boring-gate-that-fixed-it-450</link>
      <guid>https://dev.to/zubair_shaikh_a22cea567b0/i-stopped-trusting-my-ai-agents-done-heres-the-boring-gate-that-fixed-it-450</guid>
      <description>&lt;p&gt;I use AI coding agents every day for real client work. A few months ago I noticed a pattern, and once you see it you can't unsee it:&lt;/p&gt;

&lt;p&gt;The agent says "done". I push. Tests fail. Or lint breaks. Or the same wrong assumption comes back next session, because nothing was ever written down anywhere.&lt;/p&gt;

&lt;p&gt;Each instance was easy to fix. But the &lt;em&gt;pattern&lt;/em&gt; kept costing me. And the pattern wasn't a model problem — it was a verification problem. Nobody was checking the work before it left the machine.&lt;/p&gt;

&lt;h3&gt;
  
  
  The dumb fix that worked
&lt;/h3&gt;

&lt;p&gt;I stopped trying to make the agent smarter and built a gate instead. The gate is deliberately dumb — no intelligence in it at all. It's a Node CLI called &lt;strong&gt;forgekit&lt;/strong&gt; (I built it, it's MIT, zero runtime dependencies):&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. &lt;code&gt;forge verify&lt;/code&gt; — make "done" mean something checkable&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It detects the repo's test suites, runs them, and reports the truth: PASS, FAIL, INCOMPLETE, or PARTIAL. The key decision: a run that only covered part of the repo is reported as &lt;em&gt;partial&lt;/em&gt; — never as green. Most of my "done but broken" failures died right here.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8lkb5773h1sj3u4a4pgz.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8lkb5773h1sj3u4a4pgz.gif" alt="forgekit verify" width="736" height="520"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. &lt;code&gt;forge precommit&lt;/code&gt; — check before it lands&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Looks at what's actually staged and scans for secrets before anything pushes. The unglamorous stuff that saves you exactly once, and that once is worth it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbte4s7hl8u8oo09xpewx.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbte4s7hl8u8oo09xpewx.gif" alt="forgekit precommit" width="736" height="520"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. &lt;code&gt;forge impact&lt;/code&gt; — "if I touch this, what breaks?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A blast-radius estimate before an edit. Honest caveat, and I put it in the README too: it's a heuristic code graph, not a sound call graph. It misses files as well as over-warns. I treat its output as advisory. It's the part I trust least — real repos are the only way to harden a heuristic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. &lt;code&gt;forge ledger&lt;/code&gt; / &lt;code&gt;forge remember&lt;/code&gt; — memory with receipts&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Lessons and decisions live in the repo as claims, each carrying its evidence. Only tests, CI, or a human raise a claim's confidence. So the agent stops re-learning the same thing every session — and when it "remembers" something, you can see &lt;em&gt;why&lt;/em&gt; it believes it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa3zi1mpy0377dvd2mfke.gif" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa3zi1mpy0377dvd2mfke.gif" alt="forgekit ledger" width="736" height="520"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The workflow rule
&lt;/h3&gt;

&lt;p&gt;No push without &lt;code&gt;forge verify&lt;/code&gt; + &lt;code&gt;forge precommit&lt;/code&gt; green. It's a rule, not a suggestion. The repeated-failure class that motivated the whole project is gone from my workflow.&lt;/p&gt;

&lt;h3&gt;
  
  
  What I deliberately didn't build
&lt;/h3&gt;

&lt;p&gt;A sandbox — guardrails reduce risk, they're not a security boundary. And anything claiming formal verification: "proof-carrying memory" is the feature's name, not a theorem. When you're selling reliability, naming things honestly is part of the product.&lt;/p&gt;

&lt;h3&gt;
  
  
  Try it
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;npm install -g @codewithjuber/forgekit&lt;/code&gt; → &lt;code&gt;forge init&lt;/code&gt;. One init emits native config for Claude Code, Codex, Cursor, Gemini, Aider, Copilot, Windsurf, Zed, Continue and OpenClaw. (Claude Code is what I use daily, so that's the deepest-tested integration.)&lt;/p&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/CodeWithJuber/forgekit" rel="noopener noreferrer"&gt;https://github.com/CodeWithJuber/forgekit&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you try it, tell me where the impact analysis misses. That's the backlog I'm building against.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>productivity</category>
      <category>devtools</category>
    </item>
  </channel>
</rss>
