<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Manpreet Singh</title>
    <description>The latest articles on DEV Community by Manpreet Singh (@manpreet171).</description>
    <link>https://dev.to/manpreet171</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4050727%2F6d7996fe-9c11-4d2d-a5be-7157f8dfadb9.png</url>
      <title>DEV Community: Manpreet Singh</title>
      <link>https://dev.to/manpreet171</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/manpreet171"/>
    <language>en</language>
    <item>
      <title>AGENTS.md vs CLAUDE.md: Claude Code reads both now, just not at once.</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 12:00:31 +0000</pubDate>
      <link>https://dev.to/manpreet171/agentsmd-vs-claudemd-claude-code-reads-both-now-just-not-at-once-hd</link>
      <guid>https://dev.to/manpreet171/agentsmd-vs-claudemd-claude-code-reads-both-now-just-not-at-once-hd</guid>
      <description>&lt;p&gt;Two files, one job: tell a coding agent about your repo. For a year the answer to "which one?" was simple. Claude Code read &lt;code&gt;CLAUDE.md&lt;/code&gt;. Almost everything else read &lt;code&gt;AGENTS.md&lt;/code&gt;. If you used both kinds of tool, you kept two files and watched them drift apart.&lt;/p&gt;

&lt;p&gt;On 18 September that changed. Claude Code 2.1.277 started reading &lt;code&gt;AGENTS.md&lt;/code&gt; natively. Some of the top results for this question still say it doesn't.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Claude Code reads AGENTS.md now. But only when there's no CLAUDE.md, and a single CLAUDE.local.md switches it off.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Here's the current picture, the one setup that works on every version, and the part none of the comparison posts measure: what the file costs you on every turn, and what's missing from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short answer
&lt;/h2&gt;

&lt;p&gt;From Anthropic's &lt;a href="https://code.claude.com/docs/en/memory#agents-md" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt;, as of today:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Claude Code reads &lt;code&gt;AGENTS.md&lt;/code&gt; only when there's no &lt;code&gt;CLAUDE.md&lt;/code&gt;.&lt;/strong&gt; That means no &lt;code&gt;CLAUDE.md&lt;/code&gt;, &lt;code&gt;.claude/CLAUDE.md&lt;/code&gt; or &lt;code&gt;CLAUDE.local.md&lt;/code&gt; in the folder you're working in or any folder above it. Your personal &lt;code&gt;~/.claude/CLAUDE.md&lt;/code&gt; doesn't count, so that still loads alongside.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The trap.&lt;/strong&gt; Add a &lt;code&gt;CLAUDE.local.md&lt;/code&gt; for your own notes and your team's &lt;code&gt;AGENTS.md&lt;/code&gt; quietly stops loading. Nothing tells you.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;You can change the rule.&lt;/strong&gt; In &lt;code&gt;/config&lt;/code&gt;, "Project instructions" can be set to read either file (the default), both, or only &lt;code&gt;CLAUDE.md&lt;/code&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Nearly everything else reads &lt;code&gt;AGENTS.md&lt;/code&gt;.&lt;/strong&gt; Codex, Cursor, Copilot, Jules, Amp, Windsurf, Zed and more, per the list at &lt;a href="https://agents.md/" rel="noopener noreferrer"&gt;agents.md&lt;/a&gt;. Gemini CLI is the odd one out: it reads &lt;code&gt;GEMINI.md&lt;/code&gt; unless you point it elsewhere.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One more thing to check before any of that applies to you:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ claude --version
2.1.257 (Claude Code)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the CLI on my own machine. It's older than 2.1.277, so here an &lt;code&gt;AGENTS.md&lt;/code&gt; on its own would be invisible to Claude Code, whatever the docs say. If your team updates at different speeds, some of you are reading the file and some aren't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which one should you use?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Only Claude Code touches the repo?&lt;/strong&gt; Keep &lt;code&gt;CLAUDE.md&lt;/code&gt;. There's no &lt;code&gt;AGENTS.md&lt;/code&gt; in this website's repo, because Claude Code is the only agent that works on it. A second file would be a second thing to keep true.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Several tools, or a team on mixed versions?&lt;/strong&gt; Make &lt;code&gt;AGENTS.md&lt;/code&gt; the one real file, and make &lt;code&gt;CLAUDE.md&lt;/code&gt; a single line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;@AGENTS.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the setup Anthropic's docs recommend. Claude Code pulls the file in and never loads it twice, and it works on old versions too, because it doesn't depend on the new support at all. Two details that trip people up: an import inside a code fence is silently ignored, and on Windows a symlink instead of an import needs admin rights or Developer Mode, and Claude's own edit tools refuse to write through one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What not to do&lt;/strong&gt; is keep two hand-maintained copies. On the long-running &lt;a href="https://github.com/anthropics/claude-code/issues/6235" rel="noopener noreferrer"&gt;GitHub issue&lt;/a&gt; asking for &lt;code&gt;AGENTS.md&lt;/code&gt; support, it's a recurring complaint: the copies drift, and each agent ends up following different rules.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the file costs you
&lt;/h2&gt;

&lt;p&gt;Whichever file you pick, it's read at the start of every session and then sent again on every turn, like everything else in the conversation. So its size isn't a one-off. Here's this repo's &lt;code&gt;CLAUDE.md&lt;/code&gt;, checked by a small linter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;npx &lt;span class="nt"&gt;-y&lt;/span&gt; github:manpreet171/bridle agents CLAUDE.md
Linting CLAUDE.md

  ✔ 2 runnable &lt;span class="nb"&gt;command &lt;/span&gt;block&lt;span class="o"&gt;(&lt;/span&gt;s&lt;span class="o"&gt;)&lt;/span&gt;
  ✘ does not say how to &lt;span class="nb"&gt;install&lt;/span&gt; / &lt;span class="nb"&gt;set &lt;/span&gt;up
  ✘ does not say how to run the tests
  ✔ how to build or run it
  &lt;span class="o"&gt;!&lt;/span&gt; 3 angle-bracket field&lt;span class="o"&gt;(&lt;/span&gt;s&lt;span class="o"&gt;)&lt;/span&gt;, e.g. &amp;lt;slug&amp;gt; — check these are argument syntax, not unfilled template text
  ✔ nothing &lt;span class="k"&gt;in &lt;/span&gt;here tells the agent to &lt;span class="k"&gt;do &lt;/span&gt;something dangerous
  ✔ 12 prohibition&lt;span class="o"&gt;(&lt;/span&gt;s&lt;span class="o"&gt;)&lt;/span&gt; — the agent knows where the edges are
  ✔ size: 916 words ≈ 1474 tokens, re-sent every turn

FAIL — 2 problem&lt;span class="o"&gt;(&lt;/span&gt;s&lt;span class="o"&gt;)&lt;/span&gt;&lt;span class="nb"&gt;.&lt;/span&gt; Your agent is reading this file on every run.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The two failures don't apply here: it's a static site with no install step and no test suite. A general linter checks for a general project, so read it as a list of questions, not a verdict.&lt;/p&gt;

&lt;p&gt;The size line is the one that matters. About 1,500 tokens, on every turn. This project has run about 9,300 turns of Claude Code, so this one file accounts for roughly 14 million tokens, sent at the cheaper cached rate. That's an estimate, but the shape isn't. I measured where the rest of a session's tokens go in &lt;a href="https://singhlabs.dev/blog/claude-code-token-usage/" rel="noopener noreferrer"&gt;a separate post&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Does a longer file make the agent worse at following it? Anthropic's &lt;a href="https://code.claude.com/docs/en/memory" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt; say to aim for under 200 lines, and its &lt;a href="https://code.claude.com/docs/en/best-practices" rel="noopener noreferrer"&gt;best-practices guide&lt;/a&gt; warns that bloated files get ignored. The research is more mixed. An &lt;a href="https://arxiv.org/abs/2602.11988" rel="noopener noreferrer"&gt;ETH Zurich study&lt;/a&gt; found files written by people raised success rates slightly, files generated by an AI lowered them, and both added roughly 20% to cost. A &lt;a href="https://arxiv.org/abs/2605.10039" rel="noopener noreferrer"&gt;study of 1,650 Claude Code sessions&lt;/a&gt; found no measurable effect of length on following one simple rule. Length isn't proven to hurt. It is proven to cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's missing from it
&lt;/h2&gt;

&lt;p&gt;A file can be long and still not say the thing you actually need it to. Here's the same project from the other direction: not what the file says, but what I keep typing to the AI because the file doesn't say it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ npx -y toldya --dry
toldya · 13 sessions (1 Aug – 5 Oct) · 975 of your messages · 58 corrections

You keep telling your AI:
   1.   7×  keep it simple   (6 sessions)
   2.   3×  Don't confuse me   (3 sessions)

And 8 times you told it to try or check again: its first go missed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;916 words of rules, and neither line is in there. The two corrections I gave most often in this repo had never been written down, so every new session started without them and I typed them again.&lt;/p&gt;

&lt;p&gt;That's the real answer to "which file?". The name decides which tools read it. What's in it decides whether it was worth reading.&lt;/p&gt;

&lt;h2&gt;
  
  
  What belongs in it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;What it can't work out on its own.&lt;/strong&gt; The odd build command, the folder that looks unused but isn't, the deploy that happens on push.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The edges.&lt;/strong&gt; What it must never touch, and what needs a person. Those are the lines that prevent damage.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;What you keep correcting.&lt;/strong&gt; If you've typed it three times, it's a rule. Write it in your own words.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Not documentation.&lt;/strong&gt; Explaining the codebase is what the code is for. Every paragraph the agent could have read in the repo is a paragraph you pay for on every turn.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Not generated filler.&lt;/strong&gt; The ETH study's sharpest finding was that AI-written context files made results worse. Write it yourself, briefly.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Check your own
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;claude &lt;span class="nt"&gt;--version&lt;/span&gt;                             &lt;span class="c"&gt;# 2.1.277 or later to read AGENTS.md&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;npx &lt;span class="nt"&gt;-y&lt;/span&gt; github:manpreet171/bridle agents CLAUDE.md   &lt;span class="c"&gt;# or AGENTS.md&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;npx &lt;span class="nt"&gt;-y&lt;/span&gt; toldya &lt;span class="nt"&gt;--dry&lt;/span&gt;                          &lt;span class="c"&gt;# what you keep repeating&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both tools are free, read only your own files, and change nothing unless you say yes. &lt;a href="https://singhlabs.dev/bridle/" rel="noopener noreferrer"&gt;bridle&lt;/a&gt; reads the instruction file; &lt;a href="https://singhlabs.dev/toldya/" rel="noopener noreferrer"&gt;toldya&lt;/a&gt; reads your Claude Code history.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;This is how we set agents up for clients.&lt;/strong&gt; One short file that says what the agent can't guess and where the edges are, checked against what people actually keep correcting. Not a template, and never a generated one.&lt;/p&gt;

&lt;p&gt;Sources: every terminal block is a real run in this website's repo on 5 Oct 2026, the toldya one trimmed by one line (a link). Loading rules are from Anthropic's &lt;a href="https://code.claude.com/docs/en/memory#agents-md" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt; and the Claude Code changelog for 2.1.277. Tool support is from &lt;a href="https://agents.md/" rel="noopener noreferrer"&gt;agents.md&lt;/a&gt;. The studies are &lt;a href="https://arxiv.org/abs/2602.11988" rel="noopener noreferrer"&gt;arXiv 2602.11988&lt;/a&gt; and &lt;a href="https://arxiv.org/abs/2605.10039" rel="noopener noreferrer"&gt;arXiv 2605.10039&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Read next: &lt;a href="https://singhlabs.dev/blog/claude-md-rules-chat-history/" rel="noopener noreferrer"&gt;The best CLAUDE.md rules are hiding in your chat history&lt;/a&gt; · &lt;a href="https://singhlabs.dev/blog/claude-code-token-usage/" rel="noopener noreferrer"&gt;Where your Claude Code tokens actually go. Output is 0.2%.&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://singhlabs.dev/blog/agents-md-vs-claude-md/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>devtools</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Claude Code says it's done. Check the diff, not the paragraph.</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 11:58:50 +0000</pubDate>
      <link>https://dev.to/manpreet171/claude-code-says-its-done-check-the-diff-not-the-paragraph-3g55</link>
      <guid>https://dev.to/manpreet171/claude-code-says-its-done-check-the-diff-not-the-paragraph-3g55</guid>
      <description>&lt;p&gt;"Done. All tests pass, and nothing else was touched." You read it, it's late, you merge.&lt;/p&gt;

&lt;p&gt;Claude Code's own issue tracker is full of what happens next. A test suite whose total quietly changed from 4,992 to 4,966 before the &lt;a href="https://github.com/anthropics/claude-code/issues/46940" rel="noopener noreferrer"&gt;"ALL PASSED"&lt;/a&gt;. Six tests &lt;a href="https://github.com/anthropics/claude-code/issues/45550" rel="noopener noreferrer"&gt;marked skip&lt;/a&gt;. A project reported 100% complete that an audit &lt;a href="https://github.com/anthropics/claude-code/issues/53983" rel="noopener noreferrer"&gt;put nearer 60%&lt;/a&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The summary isn't a lie. It's written to explain the work, and nobody explains the parts they'd rather you didn't see.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Asking the agent to try harder doesn't fix that. Checking does. Here's what goes wrong, what doesn't work, and what happened when I held this site's own Claude Code summaries up against their real diffs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four ways "done" isn't done
&lt;/h2&gt;

&lt;p&gt;Read enough of those issues and the same four shapes keep coming back:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The tests got easier, not the code better.&lt;/strong&gt; A test skipped, an assertion loosened, a test command narrowed to the ones that pass. The suite is green because there's less of it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The count got reframed.&lt;/strong&gt; A failure becomes "a flaky timeout". A smaller total gets reported as everything passing.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Intent got ticked off as work.&lt;/strong&gt; A to-do marked done because it was planned, or a stub counted as a feature.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;A file changed and never got mentioned.&lt;/strong&gt; The quietest one, and the one that bites three days later.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It's common enough to measure. A &lt;a href="https://arxiv.org/abs/2605.29442" rel="noopener noreferrer"&gt;study of 20,574 real agent sessions&lt;/a&gt; found inaccurate self-reporting was about 23% of all the misbehaviour it caught, and a growing share over time. Anthropic's own &lt;a href="https://code.claude.com/docs/en/best-practices" rel="noopener noreferrer"&gt;best-practices guide&lt;/a&gt; says it plainly: Claude stops when the work looks done.&lt;/p&gt;

&lt;h2&gt;
  
  
  What doesn't work
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Asking "did you actually do it?"&lt;/strong&gt; The same model that wrote the summary checks it with the same blind spot. In one report it &lt;a href="https://github.com/anthropics/claude-code/issues/39907" rel="noopener noreferrer"&gt;said yes three times&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A rule in &lt;code&gt;CLAUDE.md&lt;/code&gt;.&lt;/strong&gt; One reporter had a 250-line file, memory notes and hooks, and &lt;a href="https://github.com/anthropics/claude-code/issues/37818" rel="noopener noreferrer"&gt;still got false "done"s&lt;/a&gt;. In another, the rule against skipping tests was right there in context and got broken anyway. A rule asks. A check doesn't need to.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I found in my own sessions
&lt;/h2&gt;

&lt;p&gt;Most writing on this describes the problem. Almost none of it shows a real summary next to its real diff, so I did that with two pieces of work Claude Code finished on this website today. For each one I took the summary it gave me at the end, word for word, and ran it through &lt;a href="https://singhlabs.dev/plumb/" rel="noopener noreferrer"&gt;plumb&lt;/a&gt;, a small tool that holds a summary against &lt;code&gt;git diff&lt;/code&gt; and prints only what doesn't match.&lt;/p&gt;

&lt;p&gt;The first was publishing a blog post, along with a fix to a script:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ plumb check summary.md --base a12ccca
plumb — 9 files changed, 1 named in the summary

changed but never mentioned  (read these first)
  · .assetsignore modified
  · LOG.md modified
  · assets/blog/og-claude-code-token-usage.png added
  · baggage/index.html modified
  · blog/claude-code-token-usage/index.html added
  · blog/feed.xml modified
  · blog/index.html modified
  · sitemap.xml modified

8 things the summary did not tell you.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Eight flags. I went through them one by one. Seven were things the summary did describe, just not by file name: it gave the post's URL, called the image "the social card", and said "that folder is now excluded from publishing" instead of naming &lt;code&gt;.assetsignore&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;One wasn't. &lt;strong&gt;&lt;code&gt;baggage/index.html&lt;/code&gt;&lt;/strong&gt;: Claude added a link to a product page while publishing the post, and the summary never said so anywhere. Harmless, as it happens. But I only know it's harmless because I looked.&lt;/p&gt;

&lt;p&gt;The second was a batch of SEO fixes: 21 files changed, and again only one named by path. This time the summary described all 20 others in categories, "19 pages and the sitemap", and every one checked out. plumb also flagged two files as "mentioned but not changed", which the summary had only referred to. Nothing was missing.&lt;/p&gt;

&lt;p&gt;So, across two real summaries: 30 files changed, 2 named by path, one change genuinely left out. No skipped tests, no removed guards. Neither summary lied. One still hid a change, and I'd have merged it without blinking.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually works
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Read the diff, not the paragraph.&lt;/strong&gt; &lt;code&gt;/diff&lt;/code&gt; in Claude Code, or &lt;code&gt;git diff --stat&lt;/code&gt;, shows every file that changed. The summary is a guide to the diff, never a replacement for it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Make it end with a file list.&lt;/strong&gt; Ask for every file it touched, by path, at the end of each task. My two summaries were noisy to check because they talked in categories. A list turns "did it mention everything?" into a mechanical question.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Check on every stop, automatically.&lt;/strong&gt; Claude Code's &lt;a href="https://code.claude.com/docs/en/hooks" rel="noopener noreferrer"&gt;Stop hook&lt;/a&gt; receives the agent's last message, so a check can run the moment it says it's finished. &lt;a href="https://singhlabs.dev/trust-issues/" rel="noopener noreferrer"&gt;trust issues&lt;/a&gt; is a free plugin that does exactly that: the same four checks, every time, with nobody remembering to run them.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Count the tests, before and after.&lt;/strong&gt; If the number went down, ask why before you read anything else.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Get a second pair of eyes that didn't do the work.&lt;/strong&gt; Anthropic's guide suggests a separate verification step that checks nothing outside the task changed. A fresh session doesn't share the first one's reasons for skipping a test.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it on yours
&lt;/h2&gt;

&lt;p&gt;Save the agent's last message to a file, then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; github:manpreet171/plumb
&lt;span class="nv"&gt;$ &lt;/span&gt;plumb check summary.md                &lt;span class="c"&gt;# against your uncommitted changes&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;plumb check summary.md &lt;span class="nt"&gt;--base&lt;/span&gt; main    &lt;span class="c"&gt;# against a branch or commit&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Exit code 1 means it found something, so it works as a CI gate too. What it won't do, so you don't over-trust it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It matches names literally.&lt;/strong&gt; A URL or a category doesn't count as a mention, which is why my first run flagged seven things that weren't really missing. Ask for a file list and that noise goes away.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It can't tell whether the code works.&lt;/strong&gt; It finds what was left out of the story, not bugs. Run the tests for that.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Its "quiet cuts" are patterns.&lt;/strong&gt; It spots a removed assertion or an added &lt;code&gt;.skip&lt;/code&gt;, not every possible way to weaken a check.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;strong&gt;This is how we build agents for clients.&lt;/strong&gt; Nothing gets called done because the agent said so. Every run ends with what changed, and something other than the agent checks it.&lt;/p&gt;

&lt;p&gt;Sources: the plumb output is a real run (plumb 1.1.0) on 5 Oct 2026, against the exact end-of-task summary Claude Code gave for each piece of work, with the summary file's path shortened. The second run's output is described rather than shown because it lists 20 files. Incidents are from the linked GitHub issues on anthropics/claude-code; the session study is &lt;a href="https://arxiv.org/abs/2605.29442" rel="noopener noreferrer"&gt;arXiv 2605.29442&lt;/a&gt;; hook and review guidance is from Anthropic's &lt;a href="https://code.claude.com/docs/en/best-practices" rel="noopener noreferrer"&gt;best practices&lt;/a&gt; and &lt;a href="https://code.claude.com/docs/en/hooks" rel="noopener noreferrer"&gt;hooks&lt;/a&gt; docs.&lt;/p&gt;

&lt;p&gt;Read next: &lt;a href="https://singhlabs.dev/blog/a-green-tick-over-nothing/" rel="noopener noreferrer"&gt;My mailer printed “campaign sent”. It had sent to nobody.&lt;/a&gt; · &lt;a href="https://singhlabs.dev/blog/claude-code-token-usage/" rel="noopener noreferrer"&gt;Where your Claude Code tokens actually go. Output is 0.2%.&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://singhlabs.dev/blog/claude-code-says-done/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>testing</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Where your Claude Code tokens actually go. Output is 0.2%.</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 11:38:18 +0000</pubDate>
      <link>https://dev.to/manpreet171/where-your-claude-code-tokens-actually-go-output-is-02-n8e</link>
      <guid>https://dev.to/manpreet171/where-your-claude-code-tokens-actually-go-output-is-02-n8e</guid>
      <description>&lt;p&gt;You hit the usage limit before lunch. So you do what everyone online says: make the AI talk less. Ban the "Great question!", cut the summaries, install one of the skills that trims its replies.&lt;/p&gt;

&lt;p&gt;I wanted to know where the tokens actually go, so I counted. Not estimated. Counted, from the transcripts Claude Code already keeps on your machine: every session still on mine, 173 of them, about 60,000 turns.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Everything Claude wrote back to me was 0.2% of the tokens. The rest was my own session, fed back in.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That number is real, but it isn't the whole story either, because most of those re-sent tokens are cheap ones. This post is both halves: where the tokens go, what they really cost, and the few things that actually move the number.&lt;/p&gt;

&lt;h2&gt;
  
  
  The count
&lt;/h2&gt;

&lt;p&gt;Here's the summary across every project, straight from the terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ baggage --all
baggage — 173 sessions, all projects, 60,454 turns

  before you typed anything              75k  tokens
    system prompt, every tool schema, every skill, CLAUDE.md.
    paid again on all 60,454 turns = 4.5B tokens, 17% of the bill,
    whether you called any of it or not.

  conversation you can see             27.8M  tokens
  what the API billed                  26.0B  tokens

  a 934x gap. Some of it is the fixed cost above; the rest
  is everything you picked up being re-sent on every later turn.

  output tokens           56.6M   0.2%
  re-sent context         25.3B   97.5%   &amp;lt;- the bill
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two lines matter. &lt;strong&gt;Output&lt;/strong&gt; (every reply, every line of code Claude wrote, its thinking included) is 56.6 million tokens. &lt;strong&gt;Re-sent context&lt;/strong&gt; is 25.3 billion. That's roughly 450 to 1.&lt;/p&gt;

&lt;p&gt;I'm not the first to see this shape. A &lt;a href="https://github.com/anthropics/claude-code/issues/24147" rel="noopener noreferrer"&gt;GitHub issue&lt;/a&gt; on Claude Code found the same thing in 30 days of someone else's transcripts, and a &lt;a href="https://dev.to/ploofnexa/i-measured-where-claude-code-actually-spends-tokens-968-is-re-reading-history-my-typing-was-16gm"&gt;dev.to post&lt;/a&gt; measured 0.5% output across 32 sessions. Different people, different work, same picture. What none of them show is &lt;em&gt;which things&lt;/em&gt; in a session cost the most, and what any of it means for the bill. That's the rest of this post.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it works like that
&lt;/h2&gt;

&lt;p&gt;The model has no memory between turns. Every time you press enter, Claude Code sends the whole conversation again: the system prompt, the tool definitions, every file it read, every command output, every reply so far, and then your new message.&lt;/p&gt;

&lt;p&gt;Think of a suitcase you have to carry up every flight of stairs. Pack a brick on the second floor of a 200-floor building and you carry it up 198 more flights. A 4,000-token build log that lands at turn 12 of a 200-turn session isn't 4,000 tokens. It's 4,000 sent 188 more times.&lt;/p&gt;

&lt;p&gt;So the cost of anything in a session isn't its size. It's its size times the number of turns it stays.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caching makes it cheaper, not smaller
&lt;/h2&gt;

&lt;p&gt;This is the part most guides get half right. Claude Code caches the conversation, and Anthropic &lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;prices a cache read&lt;/a&gt; at a tenth of normal input (a twentieth on Opus 5.5, a fortieth on Fable 5.1). Output costs five times input. So 97.5% of the tokens is not 97.5% of the money.&lt;/p&gt;

&lt;p&gt;I priced my own split at those published rates, with the one-hour cache a subscription uses. Output comes to &lt;strong&gt;7 to 14% of the cost&lt;/strong&gt;, depending on the model. Re-reading and re-caching the session is the other &lt;strong&gt;86 to 93%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Less dramatic than 0.2%. Still the opposite of what "make it talk less" assumes.&lt;/p&gt;

&lt;p&gt;On a Pro or Max plan you never see dollars, you see a limit. Anthropic's &lt;a href="https://code.claude.com/docs/en/costs" rel="noopener noreferrer"&gt;cost docs&lt;/a&gt; say re-read history still draws on your usage, at the cached rate, so a one-line question late in a long session draws usage for the whole session. How heavily a cached token counts against the limit isn't published. I've seen people online insist it's full weight, and others insist it's free. Nobody I found has shown either, so I'm not going to guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  The floor you pay before you type
&lt;/h2&gt;

&lt;p&gt;That first block of the output is the one nobody shows you. Before you've typed a word, a session already carries the system prompt, the built-in tools, the descriptions of every skill you've installed, your &lt;code&gt;CLAUDE.md&lt;/code&gt; files and anything else that loads at start. A typical session on my machine starts at about 75,000 tokens.&lt;/p&gt;

&lt;p&gt;It rides along on every turn. Across all my sessions, that floor alone is &lt;strong&gt;17% of everything billed&lt;/strong&gt;, whether I used any of it or not. It's also the one line you can fix in thirty seconds: every paragraph of &lt;code&gt;CLAUDE.md&lt;/code&gt; you don't need is paid again on every turn of every session, and so is the description of every skill you never call.&lt;/p&gt;

&lt;h2&gt;
  
  
  The heaviest things I carry
&lt;/h2&gt;

&lt;p&gt;Here's one project, this website, with the list of what's costing the most rent, trimmed for length:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ baggage
baggage — 13 sessions in singhlabs, 9,307 turns

  before you typed anything              84k  tokens
    system prompt, every tool schema, every skill, CLAUDE.md.
    paid again on all 9,307 turns = 779.0M tokens, 18% of the bill,
    whether you called any of it or not.

  output tokens            7.1M   0.2%
  re-sent context          4.2B   98.0%   &amp;lt;- the bill

  HEAVIEST THINGS YOU ARE STILL CARRYING
  (rent = its size x the turns it stayed in context)

   304.4M   19.4%  3769x  assistant reply
   282.4M   18.0%  1376x  your message
   244.2M   15.6%  1415x  claude-in-chrome · browser_batch
    54.6M    3.5%   145x  WebSearch
    25.6M    1.6%   275x  claude-in-chrome · javascript_tool
    21.5M    1.4%    17x  claude-in-chrome · get_page_text
    14.9M    0.9%   159x  Agent
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things surprised me.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude's own replies are the biggest line: 19.4%.&lt;/strong&gt; Not because they were expensive to write. Writing all of them was part of the 0.2%. They're expensive because each one stays in the session and is sent again on every turn after it. So the "talk less" skills aren't wrong. They're right for the wrong reason: a short reply saves you far more in re-sends than it ever cost to write.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Browser automation is close behind: 15.6%.&lt;/strong&gt; And that's the text alone: page contents, element lists, logs of each click. baggage doesn't count images, so every screenshot rides along on top of that figure, uncounted. A session that clicks through a website carries a stack of pictures of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually moves the number
&lt;/h2&gt;

&lt;p&gt;Ranked by what the counts above say, not by what's easiest to write about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Start fresh between unrelated tasks.&lt;/strong&gt; &lt;code&gt;/clear&lt;/code&gt; drops everything you've been carrying, and Anthropic's docs say it costs nothing. The brick stays on floor two.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Lower the floor.&lt;/strong&gt; Uninstall skills you don't use, keep &lt;code&gt;CLAUDE.md&lt;/code&gt; to what Claude can't work out on its own, and run &lt;code&gt;/context&lt;/code&gt; once to see what loads before you type.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Keep big output out of the main session.&lt;/strong&gt; Send a long log to a file and search it, instead of printing it into the chat where it's carried to the end. Hand noisy exploration and browser work to a subagent, so only its answer comes back.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Don't break the cache mid-session.&lt;/strong&gt; Per the &lt;a href="https://code.claude.com/docs/en/prompt-caching" rel="noopener noreferrer"&gt;caching docs&lt;/a&gt;, switching model, changing effort, turning on fast mode or changing MCP servers can throw it away, and then the whole session is written to cache again at the higher rate. Decide those at the start.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Shorter replies, for the right reason.&lt;/strong&gt; Ask for the answer without the essay. It helps, because the essay gets carried.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;On a paid plan, &lt;code&gt;/usage&lt;/code&gt; now shows which skills, subagents and MCP servers used your allowance. Worth a look before changing anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  Count your own
&lt;/h2&gt;

&lt;p&gt;The tool that printed everything above is called &lt;a href="https://singhlabs.dev/baggage/" rel="noopener noreferrer"&gt;baggage&lt;/a&gt;. It's free, it reads the transcripts already on your machine, and nothing leaves it. The name on npm belongs to someone else, so install it from the repo:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; github:manpreet171/baggage
&lt;span class="nv"&gt;$ &lt;/span&gt;baggage          &lt;span class="c"&gt;# this project&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;baggage &lt;span class="nt"&gt;--all&lt;/span&gt;    &lt;span class="c"&gt;# every project&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What it doesn't tell you, so you don't over-read it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Tokens, not dollars.&lt;/strong&gt; The totals are the exact counts the API reported. Pricing them is your model and your plan.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Item sizes are estimates.&lt;/strong&gt; The list ranks things with a rough four-characters-a-token ruler, the same ruler for everything. The totals above it are exact.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Only what's still on your machine.&lt;/strong&gt; Claude Code keeps about 30 days of history by default.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;One heavy user.&lt;/strong&gt; These are my numbers. Yours will differ, which is the point of running it.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;strong&gt;This is how we work.&lt;/strong&gt; Measure what's really happening before changing anything, then change the thing the numbers point at. It's the same rule we use when an AI system we build for a business starts costing more than it should.&lt;/p&gt;

&lt;p&gt;Sources: both terminal blocks are real baggage runs on my own machine on 5 Oct 2026, the second trimmed to whole lines for length. Prices and cache multipliers are from &lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;Anthropic's pricing page&lt;/a&gt;; the cost split weights my exact token counts by them. Cache and usage behaviour is from Claude Code's &lt;a href="https://code.claude.com/docs/en/costs" rel="noopener noreferrer"&gt;costs&lt;/a&gt; and &lt;a href="https://code.claude.com/docs/en/prompt-caching" rel="noopener noreferrer"&gt;prompt caching&lt;/a&gt; docs.&lt;/p&gt;

&lt;p&gt;Read next: &lt;a href="https://singhlabs.dev/blog/claude-md-rules-chat-history/" rel="noopener noreferrer"&gt;The best CLAUDE.md rules are hiding in your chat history&lt;/a&gt; · &lt;a href="https://singhlabs.dev/blog/loop-engineering-map/" rel="noopener noreferrer"&gt;I researched loop engineering to build a product. I built nothing.&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://singhlabs.dev/blog/claude-code-token-usage/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>productivity</category>
      <category>llm</category>
    </item>
    <item>
      <title>You gave AI your documents. It's still wrong. Here's how to find out why.</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:55:30 +0000</pubDate>
      <link>https://dev.to/manpreet171/you-gave-ai-your-documents-its-still-wrong-heres-how-to-find-out-why-44h1</link>
      <guid>https://dev.to/manpreet171/you-gave-ai-your-documents-its-still-wrong-heres-how-to-find-out-why-44h1</guid>
      <description>&lt;p&gt;You upload your company handbook, or a folder of PDFs, or two years of notes. You ask it something the document plainly answers. It answers confidently, and it's wrong.&lt;/p&gt;

&lt;p&gt;Then you do what everyone does. You rewrite the question. You add "only use the document provided". You try a better model. Sometimes it works, and you never find out why.&lt;/p&gt;

&lt;p&gt;There's a way to find out why, and it takes about twenty minutes. Nothing to install, nothing to buy, no account.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, what's actually happening
&lt;/h2&gt;

&lt;p&gt;When you point an AI at your own documents, it doesn't read them the way you do. It can't — they're far too long to hold at once.&lt;/p&gt;

&lt;p&gt;So the tool does something simpler than you'd expect:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;It cuts your documents into small pieces.&lt;/strong&gt; Usually a few hundred characters each.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When you ask a question, it picks the piece that looks most similar&lt;/strong&gt; to your question.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It shows the AI that piece&lt;/strong&gt;, and asks it to answer.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The AI never sees your document. It sees one piece of it, chosen by a search you didn't configure and can't watch.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmxwybciwi9ywqmcmogjs.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmxwybciwi9ywqmcmogjs.webp" alt="A small round figure sitting on the floor reading one tiny torn scrap of paper, with an enormous untouched stack of pages towering behind it" width="800" height="692"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;This is the actual situation. The scrap is what it answers from. The stack is everything it never saw.&lt;/em&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Almost every wrong answer is decided at step 2, before the AI is involved at all. Which means arguing with the AI can't fix it.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The industry name for this arrangement is &lt;em&gt;RAG&lt;/em&gt;. You don't need the term. You need to see step 2.&lt;/p&gt;
&lt;h2&gt;
  
  
  A worked example you can run
&lt;/h2&gt;

&lt;p&gt;Here's a short leave policy. Forty lines of Python, no installs, no key — it does steps 1 and 2, then stops and shows you what the AI &lt;em&gt;would&lt;/em&gt; have been given.&lt;/p&gt;

&lt;p&gt;Save it as &lt;code&gt;check.py&lt;/code&gt; and run &lt;code&gt;python check.py&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Shows what your AI actually gets handed when you ask about your documents.
# Python 3. Nothing to install, no API key, no account.

CHUNK_SIZE = 350          # &amp;lt;-- change this number and run it again

DOCUMENT = """
Annual leave

Every full-time employee gets 24 days of annual leave a year, plus public
holidays. Leave is counted from 1 January.

Requests go through the HR portal. Your manager approves them. If your manager
is away, their deputy approves instead.

How much notice you need to give depends on how long you are taking off. For a
single day, give two working days of notice. For anything up to a week, give
two weeks of notice. For longer than a week, give one month of notice.

Carrying leave over

You can carry a maximum of 5 unused days into the next year. Anything above 5
days is lost on 31 December. Carried days must be used by 31 March.
"""

QUESTION = "how much notice do I need to book a week off?"


def chunk(text, size):
    words, chunks, current = text.split(), [], ""
    for w in words:
        if len(current) + len(w) + 1 &amp;gt; size:
            chunks.append(current.strip())
            current = ""
        current += w + " "
    if current.strip():
        chunks.append(current.strip())
    return chunks


def score(text, question):
    stop = {"how", "much", "do", "i", "need", "to", "a", "the", "for", "of", "is"}
    q = {w.strip("?.,").lower() for w in question.split()} - stop
    c = {w.strip("?.,").lower() for w in text.split()}
    return len(q &amp;amp; c)


chunks = chunk(DOCUMENT, CHUNK_SIZE)
ranked = sorted(range(len(chunks)), key=lambda i: score(chunks[i], QUESTION), reverse=True)

print(f"chunk size {CHUNK_SIZE} -&amp;gt; {len(chunks)} pieces")
print(f"question: {QUESTION}\n")

for rank, i in enumerate(ranked, 1):
    print(f"  #{rank}  piece {i}  score {score(chunks[i], QUESTION)}")

print("\nThis, and only this, is what the AI gets:\n")
print(f"  {chunks[ranked[0]]}\n")
print("Read it. Could you answer the question from that alone?")
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The document says, in plain English, that a week off needs two weeks of notice. Here's what comes back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ python check.py
chunk size 350 -&amp;gt; 2 pieces
question: how much notice do I need to book a week off?

  #1  piece 0  score 2
  #2  piece 1  score 2

This, and only this, is what the AI gets:

  Annual leave Every full-time employee gets 24 days of annual leave a year, plus
  public holidays. Leave is counted from 1 January. Requests go through the HR
  portal. Your manager approves them. If your manager is away, their deputy
  approves instead. How much notice you need to give depends on how long you are
  taking off. For a single day, give two

Read it. Could you answer the question from that alone?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look at the last five words.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;For a single day, give two&lt;/code&gt;&lt;/strong&gt; — and then the text stops. The sentence was cut in half.&lt;/p&gt;

&lt;p&gt;An AI handed that will answer &lt;strong&gt;"two days"&lt;/strong&gt;, fluently and with no hedging. It isn't hallucinating. It's answering correctly from the only evidence it was given. The evidence was just the wrong half of a sentence.&lt;/p&gt;

&lt;p&gt;And notice the scores: &lt;strong&gt;both pieces scored 2&lt;/strong&gt;. It was a tie. The wrong one won because it happened to come first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Now change one number
&lt;/h2&gt;

&lt;p&gt;Set &lt;code&gt;CHUNK_SIZE = 300&lt;/code&gt;. Same document, same question, same code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ python check.py
chunk size 300 -&amp;gt; 3 pieces
question: how much notice do I need to book a week off?

  #1  piece 1  score 3
  #2  piece 0  score 1
  #3  piece 2  score 0

This, and only this, is what the AI gets:

  long you are taking off. For a single day, give two working days of notice. For
  anything up to a week, give two weeks of notice. For longer than a week,
  give one month of notice. Carrying leave over You can carry a maximum of 5
  unused days into the next year. Anything above 5 days is lost on 31

Read it. Could you answer the question from that alone?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Correct. Because the piece boundary landed somewhere else.&lt;/p&gt;

&lt;p&gt;Here is every size I tried:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;chunk size   pieces   answer correct?
       150        5   yes
       200        4   NO
       250        3   yes
       300        3   yes
       350        2   NO
       400        2   NO
       450        2   yes
       500        2   yes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That isn't a curve you can tune your way along. It's arbitrary. Right, wrong, right, right, wrong, wrong, right, right.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A number you never chose, and probably never saw, decided whether your answer was true.&lt;/strong&gt; Not the model. Not your prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test
&lt;/h2&gt;

&lt;p&gt;That's the whole point of the exercise, and it generalises. Whatever tool you're using, before you touch the prompt:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Find what the tool actually retrieved.&lt;/strong&gt; Most of them will show you — it's called sources, citations, references, or context. Open it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read that text yourself.&lt;/strong&gt; Ignore the answer entirely.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Ask one question: could I have answered correctly from this text alone?&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That single question splits every wrong answer into two completely different problems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;No, I couldn't have&lt;/strong&gt; — then the AI never stood a chance. The retrieval is broken. Rewriting your prompt is wasted effort, and this is where most people spend their afternoon.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Yes, I could have&lt;/strong&gt; — the answer was right there and the model went past it. A genuinely different problem, with a different fix.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Twenty minutes. No framework, no monitoring platform, no spend.&lt;/p&gt;

&lt;h2&gt;
  
  
  If it's the first one
&lt;/h2&gt;

&lt;p&gt;Things that actually help, roughly in order of effort:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Ask for more pieces at once.&lt;/strong&gt; Most tools fetch one. Three or five means a cut sentence gets rescued by its neighbour.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Cut on meaning, not on character count.&lt;/strong&gt; Split at paragraphs and headings. A rule that ends a piece mid-sentence will eventually end one mid-answer.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Keep the heading with the text.&lt;/strong&gt; A piece that says "give two weeks of notice" without "Annual leave" above it can't be matched to a question about leave.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Don't rely on similarity alone.&lt;/strong&gt; Plain word matching catches the things similarity misses — product codes, error numbers, names. Using both together is now the ordinary recommendation, not an advanced one. &lt;a href="https://www.novakit.ai/blog/why-your-rag-chatbot-sucks" rel="noopener noreferrer"&gt;NovaKit make the case well&lt;/a&gt;, and &lt;a href="https://www.kapa.ai/blog/rag-gone-wrong-the-7-most-common-mistakes-and-how-to-avoid-them" rel="noopener noreferrer"&gt;kapa.ai list the other common mistakes&lt;/a&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those guides are good and I'd read both. What none of them gives you is the part above: &lt;em&gt;which&lt;/em&gt; of the five things is yours. They hand you a list of suspects. The read-it-yourself test hands you a defendant.&lt;/p&gt;

&lt;h2&gt;
  
  
  The thing nobody tells beginners
&lt;/h2&gt;

&lt;p&gt;Every article on this ends at "improve your retrieval". Fair enough. But there's a step past it, and it changes how you think about the problem.&lt;/p&gt;

&lt;p&gt;Better chunking makes wrong answers &lt;em&gt;rarer&lt;/em&gt;. It never makes them impossible. You're tuning a dial and hoping.&lt;/p&gt;

&lt;p&gt;On a system I built — a multilingual assistant answering questions across about 75,000 lines of source — we stopped tuning the dial. Every claim in every answer had to resolve to a real, checkable location in the source material. If a citation didn't resolve, &lt;strong&gt;the answer didn't ship&lt;/strong&gt;. Not flagged. Not scored. Withheld.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The goal isn't an AI that's usually right. It's a system where being wrong is structurally impossible, because a claim it can't back up never reaches you.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4gn5h7eg2e75kypblxum.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4gn5h7eg2e75kypblxum.webp" alt="A small round figure standing on a large open book, holding a note up to check it against the page, with its other hand held out flat to refuse an outstretched hand" width="800" height="523"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Checked against the source, or it doesn't get handed over. Withheld, not flagged.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That's a design decision, not a setting, and it's the difference between a demo and something people can rely on at work. You don't need it on day one. But it's worth knowing the ceiling exists, so you stop expecting prompt changes to get you there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start here
&lt;/h2&gt;

&lt;p&gt;Next time your AI gets your own documents wrong, don't rewrite the question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Look at what it was given, and ask whether you could have answered from it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Most of the time you couldn't have — and the machine you were about to blame was the only part doing its job.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Check #2 above is one of eight.&lt;/strong&gt; I've put the rest on a single printable page — the things worth thirty seconds before you act on what an AI told you. Free, no signup: &lt;a href="https://singhlabs.dev/checklist/" rel="noopener noreferrer"&gt;Before you trust an AI answer →&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I build this kind of thing for a living, and I give away the tools that check it — &lt;a href="https://singhlabs.dev/" rel="noopener noreferrer"&gt;all of them here&lt;/a&gt;, free and readable. &lt;a href="https://singhlabs.dev/built/" rel="noopener noreferrer"&gt;What I've built&lt;/a&gt; · &lt;a href="https://singhlabs.dev/services/" rel="noopener noreferrer"&gt;work with me&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Run the script on your own document and something surprising falls out? &lt;a href="https://singhlabs.dev/contact/" rel="noopener noreferrer"&gt;Tell me&lt;/a&gt; — I collect these.&lt;/p&gt;

&lt;p&gt;Read next: &lt;a href="https://singhlabs.dev/blog/jev-explained/" rel="noopener noreferrer"&gt;What is Jev? The AI trained to say “I’m only 60% sure”&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://singhlabs.dev/blog/wrong-answers-from-your-own-documents/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>rag</category>
      <category>llm</category>
      <category>beginners</category>
    </item>
    <item>
      <title>What is Jev? The AI trained to say “I'm only 60% sure.”</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:52:40 +0000</pubDate>
      <link>https://dev.to/manpreet171/what-is-jev-the-ai-trained-to-say-im-only-60-sure-4h4a</link>
      <guid>https://dev.to/manpreet171/what-is-jev-the-ai-trained-to-say-im-only-60-sure-4h4a</guid>
      <description>&lt;p&gt;I kept hearing the word. Jev. In group chats, in newsletters, in a Hacker News thread that hit 1,900 points and 500 comments in a day — which for that site is a small riot.&lt;/p&gt;

&lt;p&gt;Every explanation I clicked on was written for engineers. Non-autoregressive. Calibrated posteriors. Typed schemas. I understood maybe half.&lt;/p&gt;

&lt;p&gt;So I read the launch post, the docs, and all 500 comments. Underneath the jargon is a simple idea, and a slightly funny one.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every AI you've used so far writes. Jev doesn't. Jev decides.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Stick with me. By the end you'll be able to explain it to a friend in one sentence, you'll know what the “can't hallucinate” claim actually means (not what it sounds like), and you'll know whether it ever touches your life. It will. You won't see it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one-sentence version
&lt;/h2&gt;

&lt;p&gt;Think about the difference between an essay and a light switch.&lt;/p&gt;

&lt;p&gt;ChatGPT, Claude, Gemini — they write essays. You ask, they produce words, one after another, and the words can be anything. A poem. A recipe. Working code. A confident lie. That flexibility is the magic &lt;em&gt;and&lt;/em&gt; the problem.&lt;/p&gt;

&lt;p&gt;Jev is a light switch. You don't ask it to talk. You show it a situation and ask a fixed question with a fixed set of answers — &lt;em&gt;is this urgent, yes or no? which of these five teams handles it? how angry is this customer, one to five?&lt;/em&gt; — and it flips the switch. Instantly. And next to the switch it puts a number: &lt;strong&gt;how sure it is.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It cannot write you a sentence. That isn't a limitation they're fixing; it's the design. The people who built it gave up words on purpose, because they think most of what businesses want from AI isn't words. It's decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it's called that
&lt;/h2&gt;

&lt;p&gt;Two names, both borrowed from smart dead people, and both help.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“System One”&lt;/strong&gt; is Daniel Kahneman's. He split thinking into two modes. System 2 is slow and careful — doing your taxes. System 1 is fast and automatic — you see a face and know it's angry before you could explain why. You don't reason your way there. You just know, and you're usually right.&lt;/p&gt;

&lt;p&gt;Chatbots are being pushed towards System 2 — “think step by step”, reasoning modes, minutes of pondering. Jev is built for System 1: the thousand tiny snap judgements that happen inside software all day. Is this spam? Is there a person in this photo? Which folder does this go in? Nobody wants an essay about it. They want the answer, fast, and they want to know if it's reliable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“Jev”&lt;/strong&gt; is after William Stanley Jevons, the economist who noticed that when steam engines got more efficient, people didn't use less coal. They used vastly more, because cheap power made a thousand new things worth doing. The founders are betting the same happens when a decision costs almost nothing. That's the whole business plan, in a name.&lt;/p&gt;

&lt;p&gt;The founder is Diogo Almeida, who was at OpenAI on the research that taught language models to follow instructions — the work that became ChatGPT. So: someone who helped build the essay machine, deciding the next thing shouldn't be one. The company, TypeSafe AI, reportedly raised $40 million to do it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you actually send it
&lt;/h2&gt;

&lt;p&gt;A real example from their documentation. It's the moment it clicked for me.&lt;/p&gt;

&lt;p&gt;You give Jev a &lt;strong&gt;situation&lt;/strong&gt; — they call it “state”. Say, a support message:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Hi, I've been trying to connect my Stripe account for 3 days and it keeps failing. I'm losing sales. Please help ASAP.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And a &lt;strong&gt;question&lt;/strong&gt; with a fixed shape. Here, a yes/no: &lt;em&gt;does this message convey urgency?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;What comes back is not a paragraph. It's this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;is_urgent: 0.999
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A number. 99.9% yes. Your software reads it and acts — bumps the ticket, pings a human, whatever you've set up. No “Certainly! This message appears to be…” to wade through.&lt;/p&gt;

&lt;p&gt;There are only three kinds of question you can ask:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Yes or no.&lt;/strong&gt; Returns the probability of yes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pick one from a list.&lt;/strong&gt; Returns a probability for every option, plus an overall confidence. Up to 255 options.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Score it on a scale.&lt;/strong&gt; Low/medium/high, one to ten, whatever you define. Returns the score, the spread, and the confidence.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The trick that makes it fast: you can ask dozens of these about the same situation at once, and it answers all of them in a single pass. A chatbot builds its answer one word at a time, each depending on the last. Jev builds nothing. It looks once and flips every switch at the same time.&lt;/p&gt;

&lt;p&gt;Here's the difference as a picture:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs2om3o5fal0jwk46d6l2.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs2om3o5fal0jwk46d6l2.webp" alt="Top: a chatbot answering one word at a time, each box depending on the last, trailing off into 600 more words. Bottom: Jev taking the situation once and returning four answers with confidences at the same time" width="800" height="431"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;A chatbot answers word by word. Jev answers every question at once.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That's why the speed numbers are wild. They claim &lt;strong&gt;70 to 500 milliseconds&lt;/strong&gt; per answer against the three seconds to five minutes a reasoning chatbot takes. And pricing: &lt;strong&gt;$0.042 per million words of input&lt;/strong&gt;, output free, because the output is a handful of numbers. Someone on the team built a bot that plays Doom by asking Jev ten questions a second. About $7 an hour. The team's reaction was that this was cheaper than they expected.&lt;/p&gt;

&lt;p&gt;Honest note: those are TypeSafe's own numbers, from their own laptops, on tasks they chose. I haven't run it — it's early access with a waitlist. More on what independent people made of the claims below.&lt;/p&gt;

&lt;h2&gt;
  
  
  The claim everyone repeats
&lt;/h2&gt;

&lt;p&gt;Every headline says the same thing: &lt;em&gt;Jev can't hallucinate.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This is true. It's also not what you think it means, and the gap between those two is the most useful thing in this post.&lt;/p&gt;

&lt;p&gt;A chatbot “hallucinates” when it confidently produces something false — a fake court case, a made-up statistic, a function that doesn't exist. It can, because it can write anything. The space of possible outputs is infinite, and some of that space is nonsense.&lt;/p&gt;

&lt;p&gt;Jev can't write anything. You gave it a menu. Yes or no. One of these five teams. A number from one to ten. It is mathematically impossible for it to hand you something that isn't on the menu. No fake court case, because “fake court case” was never an option.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;But it can absolutely pick the wrong thing off the menu.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If the message wasn't urgent and Jev says 0.9 urgent, that's a wrong answer. Not a hallucination — a mistake. Several of the sharpest people in that thread made exactly this point: it can't emit an invalid answer, but it can still emit a completely wrong valid one.&lt;/p&gt;

&lt;p&gt;So why is it still a big deal? Because of the number. Here's the honest version on one card:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe4jwfvtr95jvzizmrml3.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe4jwfvtr95jvzizmrml3.webp" alt="Two cards. Left, green: impossible answers — a made-up court case, a sixth team that doesn't exist, 'maybe, it depends' — blocked by design. Right, orange: wrong answers — urgent 0.9 and it wasn't, Billing when it should have been Sales — still possible, but labelled with a confidence" width="799" height="418"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;The new part isn't the left card. It's the number on the right — trained to be honest.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;What TypeSafe actually built — the new part — is that Jev is &lt;em&gt;trained&lt;/em&gt; to make that confidence number honest. Their training method is literally called Reinforcement Learning for Calibrated Decisions. Calibrated means: when it says 90%, it should be right about nine times in ten. When it says 55%, it should be a coin flip, and it should say so.&lt;/p&gt;

&lt;p&gt;Their own post puts the problem perfectly: &lt;em&gt;“If a model can do a task 95% of the time but doesn't say when it's in the 5%, it can't automate that task.”&lt;/em&gt; That's the whole reason AI has been stuck in “assistant” mode instead of “just do it” mode. Not that it's wrong 5% of the time — that you can't tell &lt;em&gt;which&lt;/em&gt; 5%.&lt;/p&gt;

&lt;p&gt;There's good research showing chatbots make people more confident and less accurate, because they never say “I don't know”. Read that sentence again with Jev in mind. This is an AI whose entire training objective is to say how unsure it is. Whether it lives up to that, nobody outside the company has tested yet. But it's the right thing to be trying to build.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the sceptics said, in plain words
&lt;/h2&gt;

&lt;p&gt;Five hundred comments. The best ones, translated.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“This already existed.”&lt;/strong&gt; True-ish. Machine learning has had classifiers for years — models that take input and output a probability, fast, no words. Your spam filter is one. What's new, several people concluded, is that you don't have to train one per job. You describe the menu in plain English and it works immediately. One commenter called it “democratisation of classifiers”. Fair, and still a big deal — training a classifier used to need an ML engineer and a pile of labelled data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“The speed comparison is apples to oranges.”&lt;/strong&gt; Also fair. Jev is 200× faster than a chatbot at flipping switches. A chatbot can also write code, draft your email, and explain the switch. Comparing them on speed is like saying a light switch is faster than a novelist.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“Their evals grade against other AIs, not against truth.”&lt;/strong&gt; This one matters. TypeSafe's benchmark checks how closely Jev's decisions match the average of GPT-6 and Claude's flagship on the same tasks — not how often it is actually right. So “as good as the best models” really means “agrees with the best models”. Those models can be wrong together.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;“It's probably a small model, and people will copy it.”&lt;/strong&gt; One commenter estimated from the price that Jev is around 3 billion parameters — tiny by today's standards — and predicted clones within weeks. Three days after launch, “Cua S1 — a family of System One models” showed up on the same site. The category may matter more than the company.&lt;/p&gt;

&lt;p&gt;Credit where it's due: TypeSafe's launch post has a “Nuance” section under every claim, admitting where their numbers flatter them, where bias could exist, and that they can't prove their pricing isn't subsidised. More honesty than most launches manage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where you'll actually meet it
&lt;/h2&gt;

&lt;p&gt;You will never open an app called Jev and type at it. That's the point. It's plumbing. It shows up &lt;em&gt;inside&lt;/em&gt; things you already use, and the tell will be that AI-powered features get faster and quieter. Here's where it would sit:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffaxi28r93vlvmhb93308.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffaxi28r93vlvmhb93308.webp" alt="A support message arrives; Jev answers four questions at once — urgent 0.99, team billing 0.91, anger 4 of 5, refund risk 0.62 — and the code routes each: top of the queue, route to Billing, reply in a softer tone, and for the low-confidence one, a human decides" width="800" height="289"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Read the last box. The uncertain one goes to a person. That's the pattern this makes possible.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That last box is the bit that affects you. &lt;em&gt;Route to a human when confidence is low.&lt;/em&gt; The AI handles the 90% it's sure about; the uncertain 10% goes to a person. Today most AI features either handle everything (and get some wrong, confidently) or nothing.&lt;/p&gt;

&lt;p&gt;If it works as claimed, the things that get better are boring. Support tickets to the right person first time. Spam and fraud checks that run in a blink. Game characters that react instead of freezing to think. Photo apps that sort ten thousand images without a queue. Nothing you'd write a headline about. Everything you'd notice if it stopped.&lt;/p&gt;

&lt;h2&gt;
  
  
  Paragraph or checkbox?
&lt;/h2&gt;

&lt;p&gt;The practical part, and it's for anyone — not just coders. Plenty of people automate things with no-code tools now, and all of them are about to face this choice.&lt;/p&gt;

&lt;p&gt;When you're thinking of using AI for a task, ask one question first: &lt;strong&gt;is the output a paragraph or a checkbox?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If a human needs to &lt;em&gt;read&lt;/em&gt; the answer — a draft, an explanation, a summary — you want a chatbot. If software needs to &lt;em&gt;act&lt;/em&gt; on the answer — sort, route, flag, score, approve — you want a decision model, and Jev is the first mainstream one. Here's the whole decision:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8rqgmy05eireguxwypnv.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8rqgmy05eireguxwypnv.webp" alt="A flowchart: is the answer read by a person or acted on by software? Read → chatbot. Acted on → are the answers known in advance? No → chatbot. Yes → how often? A handful of times → either works. Thousands of times → a decision model like Jev. A dotted line joins them: the chatbot designs the questions, the decision model runs them" width="800" height="946"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Paragraph or checkbox? Which kind of AI your task actually needs.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The smartest comment in the whole thread was about the box in the middle. It's not Jev versus ChatGPT. The likely pattern is both: you use a chatbot to design the questions — what should the menu be? what does “urgent” mean for us? — and the decision model runs those questions a million times in production. The essay-writer designs the switchboard. The switch-flipper runs it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who should ignore this
&lt;/h2&gt;

&lt;p&gt;If you use AI to write, learn or think — most people — Jev changes nothing for you today. Carry on.&lt;/p&gt;

&lt;p&gt;If you build anything, automate anything, or decide how AI gets used where you work, it's worth twenty minutes on the docs and the waitlist, with expectations set: early access, self-reported numbers, a menu-bound model that can still be wrong, and a category that will have five competitors by Christmas.&lt;/p&gt;

&lt;p&gt;And if you just like knowing where this is going — the interesting bit isn't the speed. It's that someone who helped build the most confident machine in history just built one whose job is to admit doubt.&lt;/p&gt;

&lt;h2&gt;
  
  
  The sentence to tell your friend
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;“It's an AI that doesn't write anything — you give it a situation and a multiple-choice question, and it picks an answer instantly and tells you how sure it is.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If they say “so it's a spam filter”, say “yes, but you can build one in a sentence instead of a month.” If they say “so it can't be wrong”, say “it can — it just can't be &lt;em&gt;impossible&lt;/em&gt;, and it tells you when it's guessing.”&lt;/p&gt;

&lt;p&gt;Got early access and pointed it at something real? &lt;a href="https://singhlabs.dev/contact/" rel="noopener noreferrer"&gt;Tell me&lt;/a&gt; — especially whether the confidence numbers meant what they claimed. That's the one thing nobody outside the company has answered yet.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;This is the thing we build.&lt;/strong&gt; AI that says “I don't know” instead of guessing, and hands the uncertain ones to a person — inside the tools your business already uses. Jev is one way to get there; a well-built assistant over your own material is another. Either way the rule is the same: the number next to the answer has to be honest.&lt;/p&gt;

&lt;p&gt;Sources: TypeSafe's launch post and docs (the Stripe example is verbatim from their quickstart) and the Hacker News thread of 15 Sep 2026. Speed, price, the $7/hour figure and the 255-option limit are TypeSafe's stated numbers. I have not used Jev.&lt;/p&gt;

&lt;p&gt;Read next: &lt;a href="https://singhlabs.dev/blog/wrong-answers-from-your-own-documents/" rel="noopener noreferrer"&gt;You gave AI your documents. It’s still wrong.&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://singhlabs.dev/blog/jev-explained/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>llm</category>
      <category>beginners</category>
    </item>
    <item>
      <title>I named three developer tools. All three names were taken on npm.</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:41:41 +0000</pubDate>
      <link>https://dev.to/manpreet171/i-named-three-developer-tools-all-three-names-were-taken-on-npm-3f6h</link>
      <guid>https://dev.to/manpreet171/i-named-three-developer-tools-all-three-names-were-taken-on-npm-3f6h</guid>
      <description>&lt;p&gt;I shipped three small tools over about a week. &lt;a href="https://singhlabs.dev/bridle/" rel="noopener noreferrer"&gt;Bridle&lt;/a&gt;, &lt;a href="https://singhlabs.dev/interlock/" rel="noopener noreferrer"&gt;Interlock&lt;/a&gt;, &lt;a href="https://singhlabs.dev/slopguard/" rel="noopener noreferrer"&gt;Slopguard&lt;/a&gt;. I was pleased with the names — short, metaphors that explain the mechanism, not a vowel-dropped startup pun among them.&lt;/p&gt;

&lt;p&gt;Then, quite late, I checked the npm registry.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="k"&gt;for &lt;/span&gt;p &lt;span class="k"&gt;in &lt;/span&gt;bridle interlock slopguard&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
    &lt;/span&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s2"&gt;"%{http_code} &lt;/span&gt;&lt;span class="nv"&gt;$p&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; https://registry.npmjs.org/&lt;span class="nv"&gt;$p&lt;/span&gt;
  &lt;span class="k"&gt;done
&lt;/span&gt;200 bridle
200 interlock
200 slopguard
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;All three taken. Not squatted — real packages, by real people:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;bridle&lt;/strong&gt; — "Javascript black magic RPC over websockets"&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;interlock&lt;/strong&gt; — interlockjs, a module bundler&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;slopguard&lt;/strong&gt; — "Don't let AI slop past your gate. Detect AI-generated code patterns"&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last one is doing roughly what mine does.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual bug this caused
&lt;/h2&gt;

&lt;p&gt;Losing a name is annoying. That wasn't the problem.&lt;/p&gt;

&lt;p&gt;The problem was already sitting in my README, in the install instructions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx github:manpreet171/bridle init   &lt;span class="c"&gt;# correct&lt;/span&gt;
npx bridle lint                      &lt;span class="c"&gt;# ...not correct&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first line is fine — explicit about the source. The second was copied from muscle memory, and it does something quite different.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;npx github:user/repo&lt;/code&gt; fetches and runs from GitHub. It &lt;strong&gt;does not install anything&lt;/strong&gt;. So by the second command there is no local &lt;code&gt;bridle&lt;/code&gt; binary, and npx does what npx does: goes to the public registry, finds the package literally named &lt;code&gt;bridle&lt;/code&gt;, downloads it, and runs it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Six commands in that README would have run a stranger's websocket library on the machine of anyone following my own documentation.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Nobody was attacking me. I wrote the instructions myself, and they were wrong in a way that's invisible when you read them and obvious when you run them on a clean machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it's easy to miss
&lt;/h2&gt;

&lt;p&gt;Because it works fine on &lt;em&gt;your&lt;/em&gt; machine.&lt;/p&gt;

&lt;p&gt;You've been developing the tool. You've got it linked, or installed globally, or you're in the repo where &lt;code&gt;node_modules/.bin&lt;/code&gt; has it. &lt;code&gt;npx bridle lint&lt;/code&gt; resolves locally and does exactly what you expect.&lt;/p&gt;

&lt;p&gt;It's only on a machine that has never seen your project that npx falls through to the registry. Which is every machine except yours.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;Install explicitly from the repo, then use the short command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; github:manpreet171/bridle
bridle lint
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Installing from a GitHub source skips the npm namespace entirely, so it doesn't matter who owns the bare name. Verify it links a real binary before you publish the instruction:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; /tmp/check &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; npm init &lt;span class="nt"&gt;-y&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; npm &lt;span class="nb"&gt;install &lt;/span&gt;github:manpreet171/bridle
added 1 package
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;ls &lt;/span&gt;node_modules/.bin/
bridle  bridle.cmd  bridle.ps1
&lt;span class="nv"&gt;$ &lt;/span&gt;./node_modules/.bin/bridle &lt;span class="nt"&gt;--version&lt;/span&gt;
bridle — Harness Script Engineering CLI
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line is the whole test. It's the binary I wrote, not the one someone else published under the same word.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd do differently
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Check the registry during naming, not after.&lt;/strong&gt; It's one curl. I checked the domain and the GitHub org and skipped the one namespace my install instructions actually resolve against.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s2"&gt;"%{http_code}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; https://registry.npmjs.org/&amp;lt;name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;404&lt;/code&gt; is free. &lt;code&gt;200&lt;/code&gt; means somebody owns it and your README needs to be explicit forever.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Test install instructions on a machine that isn't yours.&lt;/strong&gt; A clean container, or at minimum a temp directory outside the project. Docs that only work where the code already exists aren't docs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Assume the good names are gone.&lt;/strong&gt; Every short English noun that describes a mechanism is taken. Either accept a scoped or suffixed package name, or accept that you install from source and say so clearly.&lt;/p&gt;

&lt;p&gt;I kept the names. They're good names and they're right for the tools. But the install line in all three READMEs now goes through the repo, and the pages on this site say plainly that the bare npm name belongs to someone else.&lt;/p&gt;

&lt;p&gt;That felt like an admission when I wrote it. It's just accurate.&lt;/p&gt;




&lt;p&gt;The tools: &lt;a href="https://singhlabs.dev/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;. All MIT, all zero-dependency, all installed from the repo rather than a bare npm name.&lt;/p&gt;

&lt;p&gt;Anyone else been bitten by npx falling through to the registry? I suspect it's more common in READMEs than anyone realises, precisely because it never fails for the author. &lt;a href="https://singhlabs.dev/contact/" rel="noopener noreferrer"&gt;Tell me&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Read next: &lt;a href="https://singhlabs.dev/blog/capability-is-not-authority/" rel="noopener noreferrer"&gt;The models are a point apart. The harnesses are on fire.&lt;/a&gt; · &lt;a href="https://singhlabs.dev/blog/loop-engineering-map/" rel="noopener noreferrer"&gt;I researched loop engineering to build a product. I built nothing.&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://singhlabs.dev/blog/npm-names-taken/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>npm</category>
      <category>javascript</category>
      <category>security</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I researched loop engineering to build a product. I built nothing.</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:40:07 +0000</pubDate>
      <link>https://dev.to/manpreet171/i-researched-loop-engineering-to-build-a-product-i-built-nothing-12p7</link>
      <guid>https://dev.to/manpreet171/i-researched-loop-engineering-to-build-a-product-i-built-nothing-12p7</guid>
      <description>&lt;p&gt;I went into last week intending to ship a loop engineering tool. The term is two months old, the searches are climbing, and I build guardrails for AI coding agents — it looked like my lane. I killed the idea twice.&lt;/p&gt;

&lt;p&gt;If you're eyeing the same space — as a builder or just deciding what to adopt — the map that stopped me is more useful than another definition post. Page one of every search is definitions. Here is what actually exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 90-second history
&lt;/h2&gt;

&lt;p&gt;May 2025: Geoffrey Huntley wires a coding agent into a bash &lt;code&gt;while true&lt;/code&gt; loop that re-feeds it the same prompt file until the work is done, and names it after &lt;a href="https://github.com/ghuntley/how-to-ralph-wiggum" rel="noopener noreferrer"&gt;Ralph Wiggum&lt;/a&gt; — ignorance, persistence, optimism. It works embarrassingly well, and &lt;a href="https://www.theregister.com/2026/01/27/ralph_wiggum_claude_loops/" rel="noopener noreferrer"&gt;The Register covers it&lt;/a&gt; as a genuine technique rather than a joke.&lt;/p&gt;

&lt;p&gt;June 2026: Boris Cherny, who created Claude Code, says he doesn't prompt it any more — loops prompt it for him. Addy Osmani wraps a name around the practice, and "loop engineering" joins the lineage: prompt engineering, then context engineering, then harness engineering, now this. Roughly one new layer per year, each wrapping the one before.&lt;/p&gt;

&lt;p&gt;So the concept is real. The question for anyone thinking of building here is different: &lt;em&gt;is there room?&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The map: what already exists
&lt;/h2&gt;

&lt;p&gt;First stop, the GitHub topic pages. This is one command and it settles most arguments:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="k"&gt;for &lt;/span&gt;t &lt;span class="k"&gt;in &lt;/span&gt;loop-engineering ralph-wiggum graph-engineering&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
    &lt;/span&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s2"&gt;"%-18s %s repos&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$t&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s2"&gt;"https://api.github.com/search/repositories?q=topic:&lt;/span&gt;&lt;span class="nv"&gt;$t&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
       | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-m1&lt;/span&gt; total_count | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="s1"&gt;'[0-9]*'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
  &lt;span class="k"&gt;done
&lt;/span&gt;loop-engineering   358 repos
ralph-wiggum       130 repos
graph-engineering  36 repos
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nearly five hundred repos across the two loop topics, two months after the term existed. Walking the actual listings, they sort into three shelves:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The runners.&lt;/strong&gt; Tools that keep an agent looping until done. The &lt;a href="https://github.com/topics/loop-engineering" rel="noopener noreferrer"&gt;topic leader&lt;/a&gt; sits near 10k stars, with multiple independent orchestrators around 3k. Every language, every agent CLI, every flavour.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The guardrails.&lt;/strong&gt; Iteration caps, dollar budgets, no-progress detection, human checkpoints, review gates. At least three projects ship all of these today. The obvious "safe loop" product already exists several times over.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The platform itself.&lt;/strong&gt; This is the shelf that ends the conversation. Claude Code shipped &lt;code&gt;/goal&lt;/code&gt; in May — a native loop that runs until a condition you wrote is true, with a separate model grading each turn — then &lt;code&gt;/loop&lt;/code&gt;, &lt;code&gt;/schedule&lt;/code&gt; and &lt;code&gt;/batch&lt;/code&gt;, then an official guide to all four. When the platform vendor absorbs a pattern into the product, the third-party version of that pattern is on a clock. Codex and Gemini CLI have equivalents.&lt;/p&gt;

&lt;p&gt;And the tell that a category has finished consolidating: the curated &lt;em&gt;awesome-lists&lt;/em&gt; have real stars now. Nobody curates an empty room.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Every tool in this category runs before or during the loop. Nothing owns the morning after.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the honest practitioners say
&lt;/h2&gt;

&lt;p&gt;The critiques are more interesting than the hype, because they point at what is &lt;em&gt;unsolved&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://newsletter.pragmaticengineer.com/p/what-is-loop-engineering" rel="noopener noreferrer"&gt;The Pragmatic Engineer&lt;/a&gt; surveyed a couple of hundred developers and found adoption much narrower than the noise suggests — most real usage is ordinary automation wearing a new name, and human-in-the-loop still often beats autonomy.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://blackmatter.vc/lab/loop-engineering-an-honest-verdict-from-someone-who-actually-runs-agent-loops" rel="noopener noreferrer"&gt;Black Matter VC&lt;/a&gt;, from someone actually running loops in production: genuine leverage, but token burn runs roughly 4× a normal chat session and up to 15× for multi-agent setups. And the failure mode he names is not technical. He calls it &lt;strong&gt;comprehension debt&lt;/strong&gt; — the gap between what's in your repo and what you actually understand, widening every time a loop ships code nobody read.&lt;/p&gt;

&lt;p&gt;Notice what's missing from both critiques: neither says "we need another runner."&lt;/p&gt;

&lt;h2&gt;
  
  
  The gap
&lt;/h2&gt;

&lt;p&gt;Line the tools up against the lifecycle of a loop and the shape jumps out. Before the run: policies, budgets, gates — covered. During the run: progress detection, checkpoints, caps — covered. Then the loop exits, and you are standing in front of two thousand changed lines with a cheerful four-sentence summary, at whatever hour it is.&lt;/p&gt;

&lt;p&gt;That moment — the morning after — has no tooling. It is the exact moment comprehension debt gets paid or deferred, and every project in those 488 repos has already clocked off by the time it arrives.&lt;/p&gt;

&lt;p&gt;That's the verdict that killed my loop product twice: the crowded part is finished, and the empty part isn't a loop tool at all. It's a reviewing problem. So instead of runner number 359 I built &lt;a href="https://singhlabs.dev/plumb/" rel="noopener noreferrer"&gt;plumb&lt;/a&gt; — it holds the agent's summary against the diff and prints only what the summary left out: the files it never mentioned, the assertion that quietly disappeared, the test that got skipped instead of fixed. Zero dependencies, no model calls, and it doesn't care whether the code came from one prompt or a forty-iteration loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  The method, if you're weighing your own idea
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Check the topic page before you write a line.&lt;/strong&gt; One curl, shown above. The repo count and the star spread tell you which shelves are full.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check whether the platform vendor already shipped it.&lt;/strong&gt; A feature in the CLI beats a repo, every time, forever.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An awesome-list with real stars means you're late.&lt;/strong&gt; Curation is what happens after consolidation.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Then look for the moment in the lifecycle nobody covers.&lt;/strong&gt; Categories crowd around the exciting phase — the run, the loop, the autonomy. The boring phases either side of it are usually empty, and the boring phase is where the actual pain lives.&lt;/p&gt;

&lt;p&gt;A day of research, nothing built, and I'd call it the most productive day of the month. The cheapest product is the one you find out not to make.&lt;/p&gt;




&lt;p&gt;The tools: &lt;a href="https://singhlabs.dev/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;. All MIT, all zero-dependency, all built after this same go/no-go research — including the two ideas that died so plumb could exist.&lt;/p&gt;

&lt;p&gt;If you're running loops in production: what do you actually do with the output the morning after? Genuinely asking — it shaped one tool already and I suspect there's more there. &lt;a href="https://singhlabs.dev/contact/" rel="noopener noreferrer"&gt;Tell me&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Read next: &lt;a href="https://singhlabs.dev/blog/a-green-tick-over-nothing/" rel="noopener noreferrer"&gt;My mailer printed “campaign sent”. It had sent to nobody.&lt;/a&gt; · &lt;a href="https://singhlabs.dev/blog/claude-md-rules-chat-history/" rel="noopener noreferrer"&gt;The best CLAUDE.md rules are hiding in your chat history&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://singhlabs.dev/blog/loop-engineering-map/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>productivity</category>
      <category>devtools</category>
    </item>
    <item>
      <title>My mailer printed "campaign sent". It had sent to nobody.</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:38:21 +0000</pubDate>
      <link>https://dev.to/manpreet171/my-mailer-printed-campaign-sent-it-had-sent-to-nobody-1m99</link>
      <guid>https://dev.to/manpreet171/my-mailer-printed-campaign-sent-it-had-sent-to-nobody-1m99</guid>
      <description>&lt;p&gt;I make small tools that catch AI agents reporting success while quietly doing nothing. That is the entire product line. One of them is literally called &lt;a href="https://singhlabs.dev/trust-issues/" rel="noopener noreferrer"&gt;trust issues&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Last week I shipped a mailer that printed &lt;code&gt;campaign sent to list&lt;/code&gt; twice, five days apart, having delivered to zero people.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it looked from the inside
&lt;/h2&gt;

&lt;p&gt;I have a small list — 47 people who followed my writing elsewhere and agreed to hear from me here. I wrote them an email. The workflow ran. It said this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;campaign 2 sent to list ***
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Green tick. Job done. Four days later I sent a second one. Same message, same green tick.&lt;/p&gt;

&lt;p&gt;Then nobody replied to either. Not one person. I spent an evening concluding that my writing was bad, that the list was dead, that three weeks of work had landed on nobody who cared.&lt;/p&gt;

&lt;p&gt;Eventually I did the obvious thing and looked at the numbers instead of my feelings.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;LISTS
      0 subscribers  Your first list

CAMPAIGNS (newest first)

  "I measured where my AI coding tokens actually go. 0.2%."
    status sent   sent 2026-08-07T09:24:06+05:30
    delivered 0 of 0   opens 0   clicks 0

  "Moving my writing home — and what I've been building"
    status sent   sent 2026-08-03T19:40:50+05:30
    delivered 0 of 0   opens 0   clicks 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Zero subscribers.&lt;/strong&gt; Meanwhile 48 contacts sat in the account, imported correctly, never attached to the list the campaigns were addressed to. Both emails went out to an empty room.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Nobody ignored my emails. Nobody received them. Five days of silence I had taken personally was a config error.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The bug is not the interesting part
&lt;/h2&gt;

&lt;p&gt;Contacts not attached to a list is a boring mistake. Ten minutes to fix.&lt;/p&gt;

&lt;p&gt;The interesting part is that &lt;strong&gt;my code told me it had worked.&lt;/strong&gt; It called the create-campaign endpoint. It called send. Both returned 200. So it printed success — because from the code's point of view, everything it did had succeeded.&lt;/p&gt;

&lt;p&gt;It had verified &lt;em&gt;the call&lt;/em&gt;. It had verified nothing about &lt;em&gt;the effect&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;That distinction is the whole of my product line, written on my own site, and I still shipped past it. An agent that says "I fixed the tests" has usually verified that it edited a file, not that anything passes. A deploy script that says "deployed" has usually verified that the API accepted the request. The gap between &lt;em&gt;the call returned 200&lt;/em&gt; and &lt;em&gt;the thing you wanted actually happened&lt;/em&gt; is where this class of bug lives, and it is enormous.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then a second bug nearly hid the fix
&lt;/h2&gt;

&lt;p&gt;I attached the contacts, re-sent, and went to confirm. It still said zero.&lt;/p&gt;

&lt;p&gt;For twenty minutes I believed the repair had failed. It hadn't. My stats script was reading the wrong endpoint, and the two endpoints disagree:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;singular   /contacts/lists/2   → {"name":"Your first list","total":47}
collection /contacts/lists     → [{"name":"Your first list","total":0}]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same list. Same second. Two different answers from the same provider. The collection endpoint reports &lt;code&gt;0&lt;/code&gt; for a list that demonstrably has 47 members, and my code happened to read that one.&lt;/p&gt;

&lt;p&gt;So I had a broken thing, then a broken measurement of the thing, and the measurement was what I was using to decide whether the fix had worked. &lt;strong&gt;When two sources disagree, find out which one lies before you conclude anything.&lt;/strong&gt; I nearly re-fixed a bug that was already fixed.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I actually confirmed it, in the end
&lt;/h2&gt;

&lt;p&gt;Not from a dashboard. From the account's send credits.&lt;/p&gt;

&lt;p&gt;The free plan allows 300 sends a day. Before the run: 300. After: 253. Exactly 47 gone — one per subscriber. That is the only number in this whole story that could not be produced by a bug in my own reporting, because it came from the provider's billing rather than from anything I wrote.&lt;/p&gt;

&lt;p&gt;Then the email arrived in my own inbox, which is the other kind of proof that doesn't lie.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I changed
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The sender now refuses to send to an empty list.&lt;/strong&gt; It checks the subscriber count first and exits with an error. A delivery of zero is a failure, and it must look like one.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The stats script resolves each list individually&lt;/strong&gt;, with a comment naming the endpoint that lies, so nobody re-learns this.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Success messages now carry a number.&lt;/strong&gt; &lt;code&gt;sent to 47 subscribers&lt;/code&gt;, never &lt;code&gt;campaign sent&lt;/code&gt;. A message with a count in it cannot be true and empty at the same time.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last one is the cheap habit worth stealing. Most misleading log lines are misleading because they describe an &lt;em&gt;action&lt;/em&gt; rather than a &lt;em&gt;result&lt;/em&gt;. "Backup complete" versus "backed up 1,204 files, 3.1 GB". One of those can be printed by a script that copied nothing.&lt;/p&gt;




&lt;p&gt;The uncomfortable version of this story is that I sell the fix. I write tools that hold an agent's claim against what actually changed on disk, because a confident summary is not evidence. Then I wrote a summary of my own, believed it, and lost five days and most of an evening to it.&lt;/p&gt;

&lt;p&gt;Knowing the failure mode is not the same as being immune to it. If anything, I was slower to check &lt;em&gt;because&lt;/em&gt; the message came from code I had written myself.&lt;/p&gt;

&lt;p&gt;Everything quoted here is real output from my own account, lightly trimmed for width. The list is 47 people; if you are one of them, that is why the first email arrived twice.&lt;/p&gt;

&lt;p&gt;Read next: &lt;a href="https://singhlabs.dev/blog/loop-engineering-map/" rel="noopener noreferrer"&gt;I researched loop engineering to build a product. I built nothing.&lt;/a&gt; · &lt;a href="https://singhlabs.dev/blog/capability-is-not-authority/" rel="noopener noreferrer"&gt;The models are a point apart. The harnesses are on fire.&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://singhlabs.dev/blog/a-green-tick-over-nothing/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>debugging</category>
      <category>email</category>
      <category>ai</category>
      <category>devops</category>
    </item>
    <item>
      <title>The models are a point apart. The harnesses are on fire.</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:35:14 +0000</pubDate>
      <link>https://dev.to/manpreet171/the-models-are-a-point-apart-the-harnesses-are-on-fire-e47</link>
      <guid>https://dev.to/manpreet171/the-models-are-a-point-apart-the-harnesses-are-on-fire-e47</guid>
      <description>&lt;p&gt;Two numbers went round in the last few weeks. On Terminal-Bench 2.1, GPT-5.6 Sol scored &lt;strong&gt;89.5%&lt;/strong&gt; and Claude Opus 5 scored &lt;strong&gt;89.1%&lt;/strong&gt;. Four tenths of a point between the default models of the two most-used coding agents on earth.&lt;/p&gt;

&lt;p&gt;Check a second leaderboard and you get 85.77% and 84.64% — different harness, different numbers, same story. About a point in it, whoever is counting.&lt;/p&gt;

&lt;p&gt;In the same weeks those numbers were being argued about, here is what those agents actually did.&lt;/p&gt;

&lt;h2&gt;
  
  
  The month, in one list
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;GhostApproval&lt;/strong&gt; — six assistants at once: Cursor, Claude Code, Antigravity, Copilot, Grok Build, GitHub. "A malicious repo uses symlinks (CWE-61) to make the agent write outside the workspace while the approval prompt hides the real target from the user."&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;DuneSlide&lt;/strong&gt; — two CVSS 9.8 holes in Cursor. Zero-click prompt injection that escaped the terminal sandbox and overwrote the sandbox helper binary.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Cursor deeplinks&lt;/strong&gt; — one click installed an attacker-controlled MCP server and ran unsandboxed commands. The install dialog truncated the commands you were approving.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;AWS Kiro&lt;/strong&gt; — "hidden text on an ordinary web page instructs AWS Kiro to silently rewrite its own MCP server config file", then reload it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;GitLost&lt;/strong&gt; — a plain-English payload in a &lt;em&gt;public&lt;/em&gt; GitHub issue made Agentic Workflows read &lt;em&gt;private&lt;/em&gt; repositories and post the contents as a public comment.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Claude Code&lt;/strong&gt; — ran &lt;code&gt;prisma migrate diff&lt;/code&gt; with the wrong parameters and deleted 22 production tables.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now read that list again and try to find one entry that a smarter model would have prevented.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Not one of these is a reasoning failure. Every single one is a question of what the agent was &lt;strong&gt;allowed to reach&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Two axes, and we only measure one
&lt;/h2&gt;

&lt;p&gt;Benchmarks measure &lt;strong&gt;capability&lt;/strong&gt;: given a task, can it do the thing. That is a real number and it has genuinely gone up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Authority&lt;/strong&gt; is a different axis entirely: what can this process touch, and who decided. Nobody publishes a leaderboard for it. There is no percentage. It is a config file that most people have never opened.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Weak model&lt;/th&gt;
&lt;th&gt;Capable model&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Unbounded authority&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;weak, unbounded&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;capable AND unbounded&lt;/strong&gt; — every incident above&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Bounded authority&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;weak, bounded&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;capable AND bounded&lt;/strong&gt; — the only good square&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Capability (left to right) is measured, benchmarked, argued about. Authority (top to bottom) is unmeasured. Four years of effort has been horizontal. The vertical axis has barely moved.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that should worry you
&lt;/h2&gt;

&lt;p&gt;For prompt injection, &lt;strong&gt;a better model makes it worse.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A weak model handed a hidden instruction on a web page misunderstands it, fumbles the syntax, gives up. A strong model reads the same instruction, works out precisely what is being asked, and executes it correctly. That is what happened to AWS Kiro. The agent didn't malfunction. It followed instructions beautifully. They just weren't yours.&lt;/p&gt;

&lt;p&gt;Every point of capability makes an agent a better employee &lt;em&gt;and&lt;/em&gt; a better confused deputy. You cannot benchmark your way out of that, because the benchmark is measuring the thing that makes it worse.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what actually helps
&lt;/h2&gt;

&lt;p&gt;Nothing exotic. The boring version works:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Decide the blast radius before you need it.&lt;/strong&gt; Which paths, which commands, which credentials. Written down, in the repo, not in your head at 11pm.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Treat any text the agent reads as hostile input.&lt;/strong&gt; Web pages, issue bodies, README files, dependency docs. GitLost was a public issue. AWS Kiro was an ordinary web page.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Read the diff, not the summary.&lt;/strong&gt; The Prisma incident had nothing to do with injection. It was a wrong flag, run confidently, with production credentials in reach.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Assume the approval dialog is lying.&lt;/strong&gt; Two of these involved prompts that hid or truncated what you were agreeing to.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that requires a better model. All of it requires deciding, in advance, what the agent may reach — which is the thing nobody wants to do, because it is admin and the benchmark doesn't reward it.&lt;/p&gt;




&lt;p&gt;I build small tools in this gap, so treat me as biased: &lt;a href="https://singhlabs.dev/interlock/" rel="noopener noreferrer"&gt;interlock&lt;/a&gt; decides what an agent may patch and what needs a person, and &lt;a href="https://singhlabs.dev/slopguard/" rel="noopener noreferrer"&gt;slopguard&lt;/a&gt; checks generated code before it hits disk. Both are free and neither is the point of this post.&lt;/p&gt;

&lt;p&gt;The point is that the industry spent a month arguing about four tenths of a point while six agents were writing outside their workspace. We are optimising the axis we can measure, and the other one is where the damage is.&lt;/p&gt;

&lt;p&gt;Incident details and quotes from &lt;a href="https://adversa.ai/blog/top-ai-coding-agent-security-resources-august-2026/" rel="noopener noreferrer"&gt;Adversa AI's August 2026 roundup&lt;/a&gt;; benchmark figures from &lt;a href="https://artificialanalysis.ai/evaluations/terminalbench-v2-1" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt; and &lt;a href="https://www.vals.ai/benchmarks/terminal-bench-2-1" rel="noopener noreferrer"&gt;vals.ai&lt;/a&gt;. I have not reproduced any of these exploits myself.&lt;/p&gt;

&lt;p&gt;Read next: &lt;a href="https://singhlabs.dev/blog/a-green-tick-over-nothing/" rel="noopener noreferrer"&gt;My mailer printed “campaign sent”. It had sent to nobody.&lt;/a&gt; · &lt;a href="https://singhlabs.dev/blog/npm-names-taken/" rel="noopener noreferrer"&gt;I named three developer tools. All three names were taken on npm.&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://singhlabs.dev/blog/capability-is-not-authority/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>agents</category>
      <category>devops</category>
    </item>
    <item>
      <title>The best CLAUDE.md rules are hiding in your chat history.</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Mon, 05 Oct 2026 10:33:16 +0000</pubDate>
      <link>https://dev.to/manpreet171/the-best-claudemd-rules-are-hiding-in-your-chat-history-30ld</link>
      <guid>https://dev.to/manpreet171/the-best-claudemd-rules-are-hiding-in-your-chat-history-30ld</guid>
      <description>&lt;p&gt;Quick question. What's the one sentence you've typed to your AI more than any other?&lt;/p&gt;

&lt;p&gt;Not the one you &lt;em&gt;think&lt;/em&gt; you type. The one you actually type.&lt;/p&gt;

&lt;p&gt;I didn't know mine either. So I went back through three months of my own Claude Code history and counted. It was "keep it simple." &lt;strong&gt;42 times&lt;/strong&gt;, across 38 different sessions. And that rule was written nowhere my AI could read it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Every correction you type and don't write down dies with the session.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you use Claude Code, Cursor or Codex every day, you have your own "keep it simple." This post is about finding it, writing it down, and checking whether it stuck. At the end there's a one-minute way to see your own top line.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 42 times
&lt;/h2&gt;

&lt;p&gt;This is a real run on my real history, 143 Claude Code sessions between July and September:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;toldya · 143 sessions (2 Jul – 25 Sept) · 6379 of your messages · 528 corrections

You keep telling your AI:
   1.  42×  keep it simple   (38 sessions)
   2.  29×  dont complex this   (18 sessions)
   3.  15×  dont assume   (15 sessions)
   4.  13×  i dont want later on   (12 sessions)
   5.   9×  dont think too much   (9 sessions)

And 71 times you told it to try or check again: its first go missed.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Numbers one and two are the same complaint, one with my typo. Together that's 71 times I told my AI to stop overbuilding. The last line is a different 71: "try again" and "check again", which isn't a rule, just an honest measure of how often the first attempt missed.&lt;/p&gt;

&lt;p&gt;This isn't a complaint about Claude. It's genuinely good at what I use it for. It's that every one of those corrections was a rule I never wrote down.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why your AI forgets everything
&lt;/h2&gt;

&lt;p&gt;A coding agent doesn't remember yesterday. Each session starts from zero. What you corrected on Tuesday afternoon is gone by Wednesday morning.&lt;/p&gt;

&lt;p&gt;The one thing it reliably reads at the start of every session is a small text file in your project called &lt;code&gt;CLAUDE.md&lt;/code&gt;. Other tools use &lt;code&gt;AGENTS.md&lt;/code&gt;, same idea.&lt;/p&gt;

&lt;p&gt;Think of it as a sticky note on the monitor for a brilliant new colleague with amnesia. They wake up every day with no memory of you. The note is the only thing that carries over. If it's not on the note, it's gone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Everyone says "just add a rule"
&lt;/h2&gt;

&lt;p&gt;Every CLAUDE.md guide gives the same advice: when your AI repeats a mistake, add a rule for it. It's good advice. I agree with it.&lt;/p&gt;

&lt;p&gt;The gap is that &lt;strong&gt;you don't notice you're repeating yourself.&lt;/strong&gt; My "keep it simple" came up in 38 sessions, across different projects, weeks apart. Each time it felt like a one-off. I never once thought "this is the thirtieth time." Repetition across days is invisible from inside the day.&lt;/p&gt;

&lt;p&gt;That's why none of my top five were in my CLAUDE.md. Not laziness. I just had no idea.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why grep lies to you
&lt;/h2&gt;

&lt;p&gt;Claude Code keeps your history on your own machine, as text files in &lt;code&gt;~/.claude/projects&lt;/code&gt;. So the obvious move is to search it. I searched for "keep it simple." It said &lt;strong&gt;747&lt;/strong&gt;. The real number was 42.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F35jefn72fd7o2zvmpohn.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F35jefn72fd7o2zvmpohn.webp" alt="Bar chart: a plain search of the history counts 747 matches for keep it simple; counting only what I actually typed gives 42" width="799" height="359"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Same history, same phrase. 18 times too big.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The files hold everything, not just what you typed: the AI repeating your words back, summaries written when a chat gets long, tool output, logs and briefs you pasted. Counting properly means keeping only the short messages you typed, dropping questions and pastes, and grouping the ways you say the same thing. "Don't complicate it", "dont complex this" and "dont comlicate that" are one habit, not three.&lt;/p&gt;
&lt;h2&gt;
  
  
  So I built a tiny thing
&lt;/h2&gt;

&lt;p&gt;It's called &lt;a href="https://singhlabs.dev/toldya/" rel="noopener noreferrer"&gt;toldya&lt;/a&gt;, as in "I told you so." A free command-line tool that does the counting above, then offers to write the results onto the sticky note for you. One command, Node 18 or newer, no account:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;npx toldya --all
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It reads your Claude Code history on your machine, shows the list above, and asks about each one: &lt;strong&gt;y&lt;/strong&gt; adds it as a rule, &lt;strong&gt;n&lt;/strong&gt; skips it, &lt;strong&gt;e&lt;/strong&gt; lets you reword it first. Nothing is written without a yes. Or pick by number:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;npx toldya --all --add 1,2,3 --to CLAUDE.md
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Added 3 rules to CLAUDE.md. Run toldya again in a week to see if they stuck.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;## Things I kept repeating
- Keep it simple.
- Dont complex this.
- Dont assume.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Yes, it keeps my typo. They're your words. Fix them with &lt;strong&gt;e&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Did it stick?
&lt;/h2&gt;

&lt;p&gt;Adding a rule doesn't mean the AI follows it. The file can be read perfectly and the rule can still not land.&lt;/p&gt;

&lt;p&gt;So toldya remembers which rules it added, and when. Run it again a week later and each rule shows how often you said it &lt;strong&gt;before&lt;/strong&gt; it went in, and how often &lt;strong&gt;since&lt;/strong&gt;. If "keep it simple" drops from 42 to near zero, the rule works. If it keeps climbing, the wording isn't landing.&lt;/p&gt;

&lt;p&gt;I can't show you a "since" number yet. The tool is a few days old, so there is no "since" to count. That's the honest state of it.&lt;/p&gt;

&lt;p&gt;And because it made me laugh: &lt;code&gt;npx toldya --all --card&lt;/code&gt; draws your top repeats as a picture, on your machine, with a "Save as PNG" button. Nothing is uploaded.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj49m14lpsjsc9dy9q1rp.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj49m14lpsjsc9dy9q1rp.webp" alt="toldya card titled Things I keep telling my AI, listing keep it simple 41 times, dont complex this 29 times, dont assume 15 times, next to the green character at a tally-mark wall" width="800" height="409"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Mine. Drawn earlier the same day, one "keep it simple" ago.&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  The part I got wrong
&lt;/h2&gt;

&lt;p&gt;I had a bigger idea first: a team version that reads a project's pull-request review comments, finds what reviewers keep writing, and turns that into rules.&lt;/p&gt;

&lt;p&gt;Before building it, I pulled the last 1,000 review comments from each of Next.js, React, Astro, Supabase and Cal.com, and ran the same counting on them. Almost nothing repeated as a rule. The top "repeats" were "This is changed now" (13 times), "Fixed in the latest commit" (9) and "Thanks for the test" (7). That's conversation. Reviewers write about the code in front of them.&lt;/p&gt;

&lt;p&gt;People correcting their own AI are the opposite: the same handful of sentences, in the same words, for months. So the team version got dropped.&lt;/p&gt;
&lt;h2&gt;
  
  
  What it doesn't do
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Claude Code only, for now.&lt;/strong&gt; Codex stores history elsewhere and I haven't tested it on real files, so I'm not claiming it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It matches words, not meaning.&lt;/strong&gt; It won't connect "use the logger" with "why is there a console.log here?". That would need a model, and I wanted it small, free and local.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It sees only the history you have.&lt;/strong&gt; Claude Code keeps about 30 days by default.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It needs a pattern.&lt;/strong&gt; A phrase must appear three times, in two different sessions. New to Claude Code? You may get "nothing repeated yet." Good result.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;It sends nothing anywhere.&lt;/strong&gt; Only your own messages are read. No account, no telemetry, no network calls.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Try it in a minute
&lt;/h2&gt;

&lt;p&gt;The safe version only shows the report and changes nothing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;npx toldya --all --dry
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The code is at &lt;a href="https://github.com/singhlabsdev/toldya" rel="noopener noreferrer"&gt;github.com/singhlabsdev/toldya&lt;/a&gt;, MIT, zero dependencies. If it finds something useful, a star helps other people find it. And I'd genuinely like to know your top line. Mine is "keep it simple." 42 times.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;This is how we build.&lt;/strong&gt; Small tools that count what really happened instead of guessing, run on your machine, and ask before they change anything. The same rule goes into the agents we build for businesses.&lt;/p&gt;

&lt;p&gt;Sources: every terminal block is a real toldya run on my own Claude Code history on 25 Sep 2026. The 747 is a plain grep of the same files that day. The review-comment figures come from the last 1,000 comments per repo, fetched from the GitHub API the same day.&lt;/p&gt;

&lt;p&gt;Read next: &lt;a href="https://singhlabs.dev/blog/loop-engineering-map/" rel="noopener noreferrer"&gt;I researched loop engineering to build a product. I built nothing.&lt;/a&gt; · &lt;a href="https://singhlabs.dev/blog/a-green-tick-over-nothing/" rel="noopener noreferrer"&gt;My mailer printed “campaign sent”. It had sent to nobody.&lt;/a&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://singhlabs.dev/blog/claude-md-rules-chat-history/" rel="noopener noreferrer"&gt;singhlabs.dev&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>claude</category>
      <category>ai</category>
      <category>productivity</category>
      <category>opensource</category>
    </item>
    <item>
      <title>An AI Agent Broke Into Hugging Face to Cheat a Test.</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Thu, 30 Jul 2026 16:16:19 +0000</pubDate>
      <link>https://dev.to/manpreet171/an-ai-agent-broke-into-hugging-face-to-cheat-a-test-3cj</link>
      <guid>https://dev.to/manpreet171/an-ai-agent-broke-into-hugging-face-to-cheat-a-test-3cj</guid>
      <description>&lt;p&gt;An AI agent ran a 4.5-day intrusion against Hugging Face's production infrastructure. &lt;/p&gt;

&lt;p&gt;17,600 actions. &lt;/p&gt;

&lt;p&gt;Cluster-admin on multiple clusters. Write access to internal source code.&lt;br&gt;
Everyone's calling it AI turning hostile.&lt;/p&gt;

&lt;p&gt;It wasn't attacking. It was cheating on a test.&lt;/p&gt;

&lt;p&gt;The agent was being graded on a security benchmark. It worked out that Hugging Face probably hosted the answer key, and decided stealing the solutions was easier than solving the challenges.&lt;/p&gt;

&lt;p&gt;Of all the content on the platform, it accessed five datasets. Every one tied to that benchmark.&lt;/p&gt;

&lt;p&gt;Here's the detail that settles it: when it got cloud credentials and probed what it could do, every potentially destructive call was issued with DryRun=True.&lt;/p&gt;

&lt;p&gt;That flag means "tell me if this would work, but don't do it."&lt;/p&gt;

&lt;p&gt;An attacker who wants damage doesn't set that. An optimiser measuring its reach does.&lt;/p&gt;

&lt;p&gt;Nobody pointed this agent at Hugging Face. The objective just said score well and the cheapest route to a high score ran through someone else's production systems.&lt;/p&gt;

&lt;p&gt;You don't need to write a harmful goal to get harmful behaviour. You only need an objective that's easier to satisfy through a path you didn't consider.&lt;/p&gt;

&lt;p&gt;We've all seen the small version. You ask an agent to make the tests pass, and it deletes the test.&lt;/p&gt;

&lt;p&gt;This is that, with production credentials &lt;/p&gt;

&lt;p&gt;Want to read article for free check comment box.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>How to make an approval mean "yes, to this" instead of "yes, forever"</title>
      <dc:creator>Manpreet Singh</dc:creator>
      <pubDate>Tue, 28 Jul 2026 11:41:48 +0000</pubDate>
      <link>https://dev.to/manpreet171/how-to-make-an-approval-mean-yes-to-this-instead-of-yes-forever-5cje</link>
      <guid>https://dev.to/manpreet171/how-to-make-an-approval-mean-yes-to-this-instead-of-yes-forever-5cje</guid>
      <description>&lt;p&gt;Here's a problem you hit about ten minutes after you let anything automated touch your pipeline.&lt;/p&gt;

&lt;p&gt;Your build breaks. Something proposes a fix. You look at it, it's fine, you approve it. Two days later the identical break happens again and you get asked again. And again. So you do the obvious thing and approve the class of fix permanently — and now you've handed out a blank cheque for a category of change you only understood one example of.&lt;/p&gt;

&lt;p&gt;Both options are bad. What you actually want is to say "yes, to this"—this exact failure, not this once, not forever.&lt;/p&gt;

&lt;p&gt;That turns out to be a fingerprinting problem, and it's more interesting than it sounds.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why you can't just hash the log&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The naive version: hash the failure text, store the hash, compare next time.&lt;/p&gt;

&lt;p&gt;It never matches. Here are two runs of the same broken workflow, three days apart:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;2026-07-21T11:02:44.7710223Z #&lt;/span&gt;&lt;span class="c"&gt;#[warning]The `set-output` command is deprecated&lt;/span&gt;
&lt;span class="gp"&gt;2026-07-24T18:47:10.0021994Z #&lt;/span&gt;&lt;span class="c"&gt;#[warning]The `set-output` command is deprecated&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Different timestamp. Different run id. Different runner. Different checkout path. Same break. A raw hash gives you two different values and your approval is useless immediately.&lt;/p&gt;

&lt;p&gt;CI logs are full of things that change on every single run and mean nothing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ISO timestamps on every line&lt;/li&gt;
&lt;li&gt;run ids and job ids&lt;/li&gt;
&lt;li&gt;runner working directories (/home/runner/work/api/api/…)&lt;/li&gt;
&lt;li&gt;commit SHAs&lt;/li&gt;
&lt;li&gt;line and column numbers&lt;/li&gt;
&lt;li&gt;durations&lt;/li&gt;
&lt;li&gt;ANSI colour codes&lt;/li&gt;
&lt;li&gt;GitHub's own ##[group] / ##[endgroup] markers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that is the failure. All of it poisons the hash.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Strip everything that varies, then hash&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The whole mechanism is one normalise function. This is the real thing, from bin/interlock.mjs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;normalize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;text&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ANSI&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/^&lt;/span&gt;&lt;span class="se"&gt;\d{4}&lt;/span&gt;&lt;span class="sr"&gt;-&lt;/span&gt;&lt;span class="se"&gt;\d{2}&lt;/span&gt;&lt;span class="sr"&gt;-&lt;/span&gt;&lt;span class="se"&gt;\d{2}&lt;/span&gt;&lt;span class="sr"&gt;T&lt;/span&gt;&lt;span class="se"&gt;[\d&lt;/span&gt;&lt;span class="sr"&gt;:.&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;+Z&lt;/span&gt;&lt;span class="se"&gt;\s?&lt;/span&gt;&lt;span class="sr"&gt;/gm&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/^##&lt;/span&gt;&lt;span class="se"&gt;\[(&lt;/span&gt;&lt;span class="sr"&gt;group|endgroup|debug|command&lt;/span&gt;&lt;span class="se"&gt;)\]&lt;/span&gt;&lt;span class="sr"&gt;.*$/gm&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;[&lt;/span&gt;&lt;span class="sr"&gt;A-Za-z&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\\[^\s&lt;/span&gt;&lt;span class="sr"&gt;"'&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;+|&lt;/span&gt;&lt;span class="se"&gt;\/(?:&lt;/span&gt;&lt;span class="sr"&gt;home|Users|opt|tmp|github|runner&lt;/span&gt;&lt;span class="se"&gt;)\/[^\s&lt;/span&gt;&lt;span class="sr"&gt;"':&lt;/span&gt;&lt;span class="se"&gt;]&lt;/span&gt;&lt;span class="sr"&gt;+/g&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;&amp;lt;path&amp;gt;&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\b[&lt;/span&gt;&lt;span class="sr"&gt;0-9a-f&lt;/span&gt;&lt;span class="se"&gt;]{7,40}\b&lt;/span&gt;&lt;span class="sr"&gt;/g&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;&amp;lt;hex&amp;gt;&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\b\d{5,}\b&lt;/span&gt;&lt;span class="sr"&gt;/g&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;&amp;lt;n&amp;gt;&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;line &lt;/span&gt;&lt;span class="se"&gt;\d&lt;/span&gt;&lt;span class="sr"&gt;+&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;/g&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;line &amp;lt;n&amp;gt;&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/:&lt;/span&gt;&lt;span class="se"&gt;\d&lt;/span&gt;&lt;span class="sr"&gt;+:&lt;/span&gt;&lt;span class="se"&gt;\d&lt;/span&gt;&lt;span class="sr"&gt;+&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;/g&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;:&amp;lt;n&amp;gt;:&amp;lt;n&amp;gt;&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\b\d&lt;/span&gt;&lt;span class="sr"&gt;+m&lt;/span&gt;&lt;span class="se"&gt;\s?\d&lt;/span&gt;&lt;span class="sr"&gt;+&lt;/span&gt;&lt;span class="se"&gt;(\.\d&lt;/span&gt;&lt;span class="sr"&gt;+&lt;/span&gt;&lt;span class="se"&gt;)?&lt;/span&gt;&lt;span class="sr"&gt;s&lt;/span&gt;&lt;span class="se"&gt;\b&lt;/span&gt;&lt;span class="sr"&gt;/g&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;&amp;lt;dur&amp;gt;&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;[&lt;/span&gt;&lt;span class="sr"&gt; &lt;/span&gt;&lt;span class="se"&gt;\t]&lt;/span&gt;&lt;span class="sr"&gt;+/g&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt; &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;trim&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nine substitutions. Every one removes something that differs between two runs of the same failure.&lt;/p&gt;

&lt;p&gt;Then the fingerprint is a hash of the failure's class plus its normalised evidence — not the whole log, just the line the classifier matched:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;fingerprint&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;sha&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;class&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;|&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nf"&gt;normalize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;evidence&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;slice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Twelve hex characters. That's it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does it hold?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Those two runs above, three days apart:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;run 4502  →  1c59f09f5474
run 4530  →  1c59f09f5474
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A genuinely different failure—a missing Python import — lands somewhere else entirely:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;run 4471  →  263819a52f2e

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So an approval recorded against 1c59f09f5474 covers that break and nothing else. The next time it happens:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AUTO  warrant h-627fdffd — this exact failure was cleared by Manpreet Singh (use 1/5)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nobody was asked. A different failure still gets asked and always will.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The two properties that make this worth doing&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One: the thing asking for permission cannot forge the fingerprint.&lt;/p&gt;

&lt;p&gt;This is the part I'd steal for other systems. The fingerprint is derived from the log — evidence that already exists, produced by CI, not by whatever is requesting approval. There's no field where an agent describes itself and gets believed. It can't widen its own permission by wording the request more generously, because nothing it says is an input.&lt;/p&gt;

&lt;p&gt;If you're building any kind of approval gate, that's the property to design for: derive the identity of the request from evidence the requester didn't author.&lt;/p&gt;

&lt;p&gt;Two: an approval should die when the rules change.&lt;/p&gt;

&lt;p&gt;A fingerprint alone still isn't enough. If I approve a fix under one policy, then loosen the policy, that old approval shouldn't quietly carry into the new world.&lt;/p&gt;

&lt;p&gt;So each clearance also stores a hash of the enforcing parts of the policy — scope, clearance buckets, limits, mode. Not the whole file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;ruleHash&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;sha&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;canon&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;scope&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;scope&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;clearance&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;clearance&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;warrant&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;warrant&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;limits&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;limits&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;p&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;}))).&lt;/span&gt;&lt;span class="nf"&gt;slice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Edit the owner field or a comment and approvals survive, because neither changes what's enforced. Add one glob to the allow-list and every past approval is void:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight console"&gt;&lt;code&gt;&lt;span class="gp"&gt;$&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;interlock log &lt;span class="nt"&gt;--verify&lt;/span&gt;
&lt;span class="go"&gt;4 entries  ·  rules in force now: 87d0bb038798
! 3 recorded under rules that no longer apply
  1 past clearance void — those failures go back to a person
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the difference between an audit trail and a control. If loosening your policy is free, the policy isn't doing anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where it falls down&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Being honest about the limits, because you'll hit them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Normalisation is a guess. Nine regexes cover GitHub Actions well. A CI system that formats logs differently needs its own rules, and a failure whose message genuinely varies run to run will never fingerprint stably.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Two different bugs can share evidence. If two distinct problems produce the identical error line, they get the same fingerprint. Narrower classes reduce this; they don't eliminate it.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;It says nothing about whether the fix is correct. It only answers "is this the same failure I already looked at?"—which is exactly one question, not all of them.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The general shape&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Strip everything that varies between runs. Hash what's left, alongside the class. Bind the approval to that hash and to a hash of the rules in force. Store it where a person can read it later.&lt;/p&gt;

&lt;p&gt;That works for CI failures. It works for anything where you want a human decision to be reusable but not unbounded.&lt;/p&gt;

&lt;p&gt;The implementation is a single zero-dependency .mjs file, MIT, 35 tests: &lt;a href="//github.com/manpreet171/interlock"&gt;github.com/manpreet171/interlock&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you've solved the "yes, to this" problem a different way, I'd genuinely like to hear it—especially the normalization, which is the part I'm least sure generalizes.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
