<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ivan</title>
    <description>The latest articles on DEV Community by Ivan (@irr123).</description>
    <link>https://dev.to/irr123</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4160629%2F3780bd10-e035-495e-8b91-ac601fff4a64.jpg</url>
      <title>DEV Community: Ivan</title>
      <link>https://dev.to/irr123</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/irr123"/>
    <language>en</language>
    <item>
      <title>Rotten specs</title>
      <dc:creator>Ivan</dc:creator>
      <pubDate>Thu, 08 Oct 2026 04:54:12 +0000</pubDate>
      <link>https://dev.to/irr123/rotten-specs-83a</link>
      <guid>https://dev.to/irr123/rotten-specs-83a</guid>
      <description>&lt;p&gt;Say spec-driven development out loud and someone answers "spec rot". OpenSpec, spec-kit and BMAD earned that answer; every one of them builds a knowledge base I never asked for.&lt;/p&gt;

&lt;p&gt;But I run SDD without rotten specs. The fix: stop building a knowledge base on top of it.&lt;/p&gt;

&lt;p&gt;SDD is two steps: write the spec, then build from it. Both steps run against a context budget, because the &lt;em&gt;1M&lt;/em&gt; window is still a marketing number.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two steps don't need a local issue tracker
&lt;/h2&gt;

&lt;p&gt;OpenSpec gave me openspec/specs/, then openspec/changes/&amp;lt;name&amp;gt;/ with proposal.md, design.md, tasks.md and a nested specs/&amp;lt;capability&amp;gt;/spec.md, then openspec/changes/archive/2025-01-23-&amp;lt;name&amp;gt;/.&lt;/p&gt;

&lt;p&gt;Matt Pocock's skills sell the opposite. Marketed as "small, easy to adapt, and composable", a shot at GSD, BMAD and Spec-Kit for "owning the process". Then /setup-matt-pocock-skills writes docs/agents/issue-tracker.md and a triage label vocabulary, /domain-modeling maintains a root CONTEXT.md plus docs/adr/0001-*.md, /to-spec publishes .scratch/&amp;lt;feature&amp;gt;/spec.md, /to-tickets fills .scratch/&amp;lt;feature&amp;gt;/issues/01-*.md, and /wayfinder lays a map.md over the pile. Smaller units that build the same knowledge base.&lt;/p&gt;

&lt;p&gt;The rest repeat the shape. A constitution, a spec, a plan and a task list per feature, a memlog. Scott Logic ran spec-kit on one feature and counted 2,577 lines of generated markdown against 689 lines of code, plus 3.5 hours of review&lt;sup id="fnref1"&gt;1&lt;/sup&gt;. I read that once. The agent reads it every session, long after the code moved on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nobody owns the memory
&lt;/h2&gt;

&lt;p&gt;Before writing this I deleted about 250MB from &lt;code&gt;~/.claude&lt;/code&gt;. Memories, plans, task lists, session leftovers, backups, security logs. Nothing broke. Nothing noticed.&lt;/p&gt;

&lt;p&gt;Count the layers that claimed to remember something for me:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the harness: projects/&amp;lt;slug&amp;gt;/memory/MEMORY.md and its per-type files, plans/, sessions/, history.jsonl, backups/

&lt;ul&gt;
&lt;li&gt;global and per project&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;the SDD framework. Constitution, spec, plan, task list, CONTEXT.md, ADRs, archive&lt;/li&gt;
&lt;li&gt;rotten AI-written comments in code, describing the shape before the last refactor&lt;/li&gt;
&lt;li&gt;mine, and the only ones I edit: &lt;a href="https://bogomolov.work/blog/posts/ai-agent-architecture-model-harness-intent/#opencode-the-open-source-ai-coding-agent" rel="noopener noreferrer"&gt;AGENTS.md&lt;/a&gt;, README.md, my tests, my openapi.yaml&lt;/li&gt;
&lt;li&gt;somebody's Jira/Monday/Confluence MCP servers

&lt;ul&gt;
&lt;li&gt;plus the default expectation that &lt;code&gt;gh&lt;/code&gt; is installed and GitHub Issues is where I track work&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A lot of places, a lot of writers. Zero owners. One deleter: me.&lt;/p&gt;

&lt;p&gt;Where to put a &lt;code&gt;{&lt;/code&gt; is a &lt;code&gt;make fmt&lt;/code&gt; question. Nobody needs to search three trackers, two memory stores and an ADR to answer it.&lt;/p&gt;

&lt;p&gt;So, just stop mixing the jobs. One tool, one job is old advice and it holds here.&lt;/p&gt;

&lt;h2&gt;
  
  
  My SDD adds no files to the repo
&lt;/h2&gt;

&lt;p&gt;I run Pi now. Zero bloatware, zero opinions I didn't add myself. My global &lt;a href="https://bogomolov.work/blog/posts/ai-agent-architecture-model-harness-intent/#opencode-the-open-source-ai-coding-agent" rel="noopener noreferrer"&gt;AGENTS.md&lt;/a&gt; is 9 lines: six invariants, two headers and one environment note. Next to it, &lt;a href="https://github.com/irr123/lean" rel="noopener noreferrer"&gt;three skills&lt;/a&gt; I wrote&lt;sup id="fnref2"&gt;2&lt;/sup&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Research&lt;/strong&gt; only reads. It splits a question into angles, sends read-only sub-agents at the independent ones, returns cited findings and writes nothing. &lt;strong&gt;Transfer&lt;/strong&gt; writes one owner-only Markdown handoff into the temp directory, secrets redacted, then stops. The handoff carries whatever the session produced: findings, or the spec. &lt;strong&gt;Apply&lt;/strong&gt; takes that handoff, checks its claims against the current repo, makes the smallest change and runs the repo's own checks. Durable guidance goes into the instruction files I already keep. No intermediate artifacts land in the repo.&lt;/p&gt;

&lt;p&gt;Past 100+k tokens in a session I do &lt;strong&gt;Transfer&lt;/strong&gt; and restart from the handoff. Same move once the spec is ready: hand off, drop the model a tier (Opus to Sonnet, Sol to Terra), then &lt;strong&gt;Apply&lt;/strong&gt;. The spec crosses one boundary and dies there.&lt;/p&gt;

&lt;p&gt;Even sessions I don't want to keep at all: &lt;code&gt;export PI_CODING_AGENT_SESSION_DIR=${TMPDIR%/}/pi-sessions&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Stop arguing with a pointless "grilling" session about &lt;em&gt;marketing strategy&lt;/em&gt; while trying to ship a local POC.&lt;/p&gt;

&lt;p&gt;While models earn more intelligence, your codebase is still the source of truth. The agent gets a new version every hour. The repo is the one worth investing in.&lt;/p&gt;

&lt;p&gt;The software development cycle has been here for decades. LLM marketing tries to bury it. Rotten specs are the punishment for blind faith in the Machine God.&lt;/p&gt;




&lt;ol&gt;

&lt;li id="fn1"&gt;
&lt;p&gt;Scott Logic, "Putting Spec Kit through its paces: radical idea or reinvented waterfall?", 2025-11-26. 2,577 lines of markdown, 689 loc, 3.5 hrs review against 1,000 loc, no markdown, 15 min review. &lt;a href="https://blog.scottlogic.com/2025/11/26/putting-spec-kit-through-its-paces-radical-idea-or-reinvented-waterfall.html" rel="noopener noreferrer"&gt;https://blog.scottlogic.com/2025/11/26/putting-spec-kit-through-its-paces-radical-idea-or-reinvented-waterfall.html&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn2"&gt;
&lt;p&gt;Not ideal. Three files, my taste, my harness. If you can do it better, don't hesitate: open a PR.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>devops</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Null hypothesis: AI code differs</title>
      <dc:creator>Ivan</dc:creator>
      <pubDate>Thu, 08 Oct 2026 04:50:29 +0000</pubDate>
      <link>https://dev.to/irr123/null-hypothesis-ai-code-differs-f65</link>
      <guid>https://dev.to/irr123/null-hypothesis-ai-code-differs-f65</guid>
      <description>&lt;p&gt;I spot AI code in PRs to repos I know. Usually it means nothing. Sometimes it annoys me. Like recognizing a colleague’s style. But a colleague learns; an agent needs an external gate saying "don't do it this way".&lt;/p&gt;

&lt;h2&gt;
  
  
  If something exists, I could measure it
&lt;/h2&gt;

&lt;p&gt;I wanted a reproducible signature to work with. Started from the observation I already had: AI code looks overengineered to me. The plan was to compute target metrics over time/commits and watch the curve bend around agent adoption.&lt;/p&gt;

&lt;p&gt;I tried cyclomatic complexity, cognitive complexity, max nesting, conditions per entity, and call stack depth.&lt;/p&gt;

&lt;p&gt;The first snag was size. Complexity tracked SLOC almost one to one. So I normalized everything on SLOC and kept looking.&lt;/p&gt;

&lt;p&gt;Then added methods and functions per entity, entities per API method, entities per package, dependency count, edges in the dependency graph, and error handling density. About 280 metrics in all with their pairwise correlations.&lt;/p&gt;

&lt;p&gt;I found no anomalies outside statistical error. Same for public repos and my private ones.&lt;/p&gt;

&lt;h3&gt;
  
  
  The explanation I couldn't shake
&lt;/h3&gt;

&lt;p&gt;As per my understanding, a model emits roughly the mean of its training distribution, while each human deviates from that mean. The means land on top of each other, and the statistic goes quiet.&lt;/p&gt;

&lt;p&gt;If that's what happened, comparing AI with human code hides the signal. I would need to compare it with one specific person instead, but I don't have enough single-author code to run that test.&lt;/p&gt;

&lt;h2&gt;
  
  
  The public papers
&lt;/h2&gt;

&lt;p&gt;The process did one useful thing: it taught me the proper questions, and those led me to people who had already run the measurements.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2603.27130" rel="noopener noreferrer"&gt;A Large-Scale Comprehensive Measurement of AI-Generated Code in Real-World Repositories&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2609.17598" rel="noopener noreferrer"&gt;Not All Agents Are Equal: Code Quality and Post-Merge Maintenance Across Five Autonomous Coding Agents in the Wild&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://arxiv.org/abs/2508.21634" rel="noopener noreferrer"&gt;Human-Written vs. AI-Generated Code: A Large-Scale Study of Defects, Vulnerabilities, and Complexity&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;These papers do find differences:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;raw&lt;/strong&gt; AI-involved code is "more comment-heavy, less cross-file reused, ..."

&lt;ul&gt;
&lt;li&gt;these differences disappear in real-world repos (my guess: &lt;a href="https://bogomolov.work/blog/posts/ai-agent-architecture-model-harness-intent/" rel="noopener noreferrer"&gt;AGENTS.md&lt;/a&gt; customizations, linters and formatters, plus review)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;The stronger signal: AI-involved commits are smaller and more localized.

&lt;ul&gt;
&lt;li&gt;which I didn't measure and don't care about&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The papers also add an angle I missed: vendors differ from each other. Which literally means there is no common signature.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Open niche:&lt;/em&gt; no one tests a single vendor against different AGENTS.md files and related tunings. My bet: another layer of tuning splits the signature again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;No general signature. Curiosity satisfied.&lt;/p&gt;

&lt;p&gt;Referring to my original annoyance, looks like there is only one way to handle it automatically: write dedicated static checkers to prevent specific patterns.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>codequality</category>
      <category>devops</category>
    </item>
    <item>
      <title>Security you can't justify is a vicious cycle</title>
      <dc:creator>Ivan</dc:creator>
      <pubDate>Thu, 08 Oct 2026 04:47:02 +0000</pubDate>
      <link>https://dev.to/irr123/security-you-cant-justify-is-a-vicious-cycle-4mpp</link>
      <guid>https://dev.to/irr123/security-you-cant-justify-is-a-vicious-cycle-4mpp</guid>
      <description>&lt;p&gt;I've already written about &lt;a href="https://bogomolov.work/blog/posts/html-sanitization/" rel="noopener noreferrer"&gt;compliance security&lt;/a&gt;. Here is the adversarial round.&lt;/p&gt;

&lt;p&gt;Dependabot continuously DDoSes me with a wall of pull requests. &lt;code&gt;npm audit&lt;/code&gt; always paints &lt;code&gt;X high-severity vulnerabilities&lt;/code&gt;. And a devops engineer, now and then, starts arguing that our microservices, inside one locked-down AWS VPC, should really talk over TLS.&lt;/p&gt;

&lt;p&gt;Different mouths, one noise: &lt;strong&gt;add more security, now.&lt;/strong&gt; Not one of them showed me an attacker. Not one showed me the threat.&lt;/p&gt;

&lt;p&gt;That's the tell. A vulnerability is worth acting on when you can show the code that exploits it. The red banner, the CVSS score, the "we should really" line: all claims. Mostly marketing. Security you can't justify becomes theater. Then it becomes a vicious cycle.&lt;/p&gt;

&lt;h2&gt;
  
  
  Talk is cheap. Show me the code.&lt;sup id="fnref1"&gt;1&lt;/sup&gt;
&lt;/h2&gt;

&lt;p&gt;A working proof of concept is proof. A severity score is a claim. A scanner label is a claim. Someone's opinion is a claim. Claims are cheap. You can generate a thousand a day, no code required. The PoC is the tax that separates the real from the loud.&lt;/p&gt;

&lt;p&gt;Watch a claim meet the tax. A hyped model scanned curl and reported five &lt;em&gt;confirmed&lt;/em&gt; vulnerabilities; after the security team looked, one was real.&lt;sup id="fnref2"&gt;2&lt;/sup&gt; The label said confirmed. The code said otherwise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Never update dependencies unless it breaks for your users.&lt;sup id="fnref3"&gt;3&lt;/sup&gt;
&lt;/h2&gt;

&lt;p&gt;Brainlessly following automated-tool guidance is a security risk. Brainlessly adding a dependency to your code or infra is a security risk.&lt;/p&gt;

&lt;p&gt;Supply chain: attackers force-pushed malicious commits to 75 of 76 tags of the &lt;em&gt;official&lt;/em&gt; Trivy Action and turned it into a credential stealer.&lt;sup id="fnref4"&gt;4&lt;/sup&gt; That proof never came from the red banner.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;None of this is anti-tooling. I keep the scanner and the model. The failure is treating a claim as a verdict: bad tool UI, plus blind obedience to decade-old compliance checklists.&lt;/p&gt;

&lt;p&gt;So I've started to understand the projects that refuse AI contributions.&lt;/p&gt;




&lt;ol&gt;

&lt;li id="fn1"&gt;
&lt;p&gt;Linus Torvalds, linux-kernel mailing list, 25 August 2000. &lt;a href="http://lkml.org/lkml/2000/8/25/132" rel="noopener noreferrer"&gt;http://lkml.org/lkml/2000/8/25/132&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn2"&gt;
&lt;p&gt;Daniel Stenberg, "Mythos finds a curl vulnerability", daniel.haxx.se, 2026-05-11. &lt;a href="https://daniel.haxx.se/blog/2026/05/11/mythos-finds-a-curl-vulnerability/" rel="noopener noreferrer"&gt;https://daniel.haxx.se/blog/2026/05/11/mythos-finds-a-curl-vulnerability/&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn3"&gt;
&lt;p&gt;Mitchell Hashimoto: "Fork your dependencies, trim them to only your use case, never update unless it breaks for your users." &lt;a href="https://x.com/mitchellh/status/2057171518027887035" rel="noopener noreferrer"&gt;https://x.com/mitchellh/status/2057171518027887035&lt;/a&gt;.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn4"&gt;
&lt;p&gt;Stephen Thoemmes, "Trivy GitHub Actions Supply Chain Compromise", Snyk. &lt;a href="https://snyk.io/articles/trivy-github-actions-supply-chain-compromise/" rel="noopener noreferrer"&gt;https://snyk.io/articles/trivy-github-actions-supply-chain-compromise/&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;/ol&gt;

</description>
      <category>security</category>
      <category>devops</category>
      <category>ai</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Runtime cost matters to me now</title>
      <dc:creator>Ivan</dc:creator>
      <pubDate>Sun, 04 Oct 2026 20:13:15 +0000</pubDate>
      <link>https://dev.to/irr123/runtime-cost-matters-to-me-now-4mdm</link>
      <guid>https://dev.to/irr123/runtime-cost-matters-to-me-now-4mdm</guid>
      <description>&lt;p&gt;AI made me reopen a boring question: why am I still paying Fargate rent for V8?&lt;/p&gt;

&lt;p&gt;Bun's Zig-to-Rust AI rewrite got me running the math. 6 days, 960'000 lines, 99.8% of Bun's existing tests passing, 9 days to merge ($10'000 total AI spent?)&lt;sup id="fnref1"&gt;1&lt;/sup&gt;.&lt;/p&gt;

&lt;p&gt;What does that look like for 15'844 LOC of my own microservice on Fargate?&lt;/p&gt;

&lt;h2&gt;
  
  
  My setup
&lt;/h2&gt;

&lt;p&gt;Node.js microservice, talks to MongoDB, calls a few neighbors, responds to others, does its small thing. A typical service sits at 0.5 vCPU / 1 GB RAM: &lt;em&gt;$18/month&lt;/em&gt; per task. Three replicas minimum for robustness, so &lt;em&gt;$54/month&lt;/em&gt; per service.&lt;/p&gt;

&lt;p&gt;Most of that bill is provisioned runtime headroom. Node makes that headroom harder to shrink. V8, dynamic dispatch, garbage collection, startup time, memory spikes. The runtime pays rent on Fargate while I'm paying for it.&lt;/p&gt;

&lt;h3&gt;
  
  
  One rewrite, run the math
&lt;/h3&gt;

&lt;p&gt;Say &lt;em&gt;$100&lt;/em&gt; in &lt;a href="https://bogomolov.work/blog/posts/ai-agent-architecture-model-harness-intent/#claude-code" rel="noopener noreferrer"&gt;Claude Code&lt;/a&gt; tokens to produce a first port to Go or Rust. Not to ship it blindly. Review, tests, and rollout still cost human time. But the first draft is no longer the expensive part.&lt;/p&gt;

&lt;p&gt;Fargate bills &lt;em&gt;$0.04048&lt;/em&gt; per vCPU-hour and &lt;em&gt;$0.004445&lt;/em&gt; per GB-hour&lt;sup id="fnref2"&gt;2&lt;/sup&gt;. Minimum task size: 0.25 vCPU / 0.5 GB, about &lt;em&gt;$9/month&lt;/em&gt;. A compiled service with a Mongo driver and a couple of HTTP clients belongs at that floor. No V8. No managed runtime heap sitting around waiting for traffic.&lt;/p&gt;

&lt;p&gt;New bill for that service: 3 x $9 = &lt;em&gt;$27/month&lt;/em&gt;; &lt;strong&gt;2x savings&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Payback on &lt;em&gt;$100&lt;/em&gt; of AI: under four months; for one service in minimal setup.&lt;/p&gt;

&lt;p&gt;That is the toy version. My real setup has dev, prod, workload replicas, autotests, and the usual operational mess. Around ten replicas. Payback lands near one month.&lt;/p&gt;

&lt;h3&gt;
  
  
  The cluster framing
&lt;/h3&gt;

&lt;p&gt;The language map changes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Rust / C / C++&lt;/strong&gt;: smallest runtime footprint. Highest migration friction. Ecosystem question still open for my services.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Go&lt;/strong&gt;: small runtime tax. Boring deployment, strong backend ecosystem. Looks like the new default from the cost side.&lt;sup id="fnref3"&gt;3&lt;/sup&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Node / Python / PHP / Java / C#&lt;/strong&gt;: still productive defaults, but their runtime footprint is no longer free. Fargate turns it into a line item.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The runtime tax used to hide inside developer productivity. AI lowers the cost of moving code. Cloud exposes the cost of keeping code running. That makes runtime choice visible again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Plan B
&lt;/h2&gt;

&lt;p&gt;Or move off AWS to &lt;a href="https://bogomolov.work/blog/posts/the-actual-state-of-self-hosting-on-a-vps/" rel="noopener noreferrer"&gt;Hetzner&lt;/a&gt; and cut the bill by ~10x.&lt;/p&gt;




&lt;ol&gt;

&lt;li id="fn1"&gt;
&lt;p&gt;&lt;a href="https://www.theregister.com/2026/05/14/anthropics_bun_rust_rewrite_merged_at_speed_of_ai/5240381" rel="noopener noreferrer"&gt;https://www.theregister.com/2026/05/14/anthropics_bun_rust_rewrite_merged_at_speed_of_ai/5240381&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn2"&gt;
&lt;p&gt;&lt;a href="https://aws.amazon.com/fargate/pricing/" rel="noopener noreferrer"&gt;https://aws.amazon.com/fargate/pricing/&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn3"&gt;
&lt;p&gt;I’ve personally never been a Rust fan. Go, though, I’ve used for heavily loaded things. C/C++ lives somewhere in old education.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>aws</category>
      <category>go</category>
      <category>devops</category>
    </item>
    <item>
      <title>AI agent architecture: model, harness and intent</title>
      <dc:creator>Ivan</dc:creator>
      <pubDate>Sat, 03 Oct 2026 23:31:16 +0000</pubDate>
      <link>https://dev.to/irr123/ai-agent-architecture-model-harness-and-intent-3018</link>
      <guid>https://dev.to/irr123/ai-agent-architecture-model-harness-and-intent-3018</guid>
      <description>&lt;p&gt;My VPS runs a "personal AI agent". It forgets its own abilities every morning.&lt;/p&gt;

&lt;p&gt;My terminal runs a coding agent. It ships production work.&lt;/p&gt;

&lt;p&gt;Same year. Frontier models on both. Same ecosystem.&lt;/p&gt;

&lt;p&gt;Both are model + harness, trying to handle the same thing: my intent. Why such a&lt;br&gt;
different experience? Start with the thing everyone mixes up: definitions.&lt;/p&gt;
&lt;h2&gt;
  
  
  What is an AI agent?
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Anthropic&lt;/strong&gt;: the company. Founded by ex-OpenAI researchers. Usual story.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude&lt;/strong&gt;: the model family. When someone says "ask Claude," they usually
mean a single call to whichever model through a chatbot. Competes with GPT
(OpenAI) and Gemini (Google).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Opus / Sonnet / Haiku&lt;/strong&gt;: model tiers. Opus = most capable, Haiku = fastest /
cheapest.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agentic approach&lt;/strong&gt; / &lt;strong&gt;Agentic AI&lt;/strong&gt;&lt;sup id="fnref1"&gt;1&lt;/sup&gt;: reason → act → observe → repeat,
the loop itself; &lt;em&gt;the loop is what makes something agentic&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent&lt;/strong&gt;: model wrapped in a harness, spins in a loop. Reads context, plans,
acts, verifies, repeats. Distinct from a chatbot by execution pattern: a chat
interface can front an agent; a one-shot call cannot.

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sub-Agent&lt;/strong&gt;: execution thread inside an agent. Inherits base settings,
extends them.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent harness&lt;/strong&gt;&lt;sup id="fnref2"&gt;2&lt;/sup&gt;: everything in an AI agent except the model itself.
Tools, memory, workflow (the loop, plan/build sub-agents), guardrails
(permissions, sandbox).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MCP server&lt;/strong&gt;: tool through which the agent interacts with the outer world.
Web, DBs, clouds, apps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Skill&lt;/strong&gt;: goal-aimed prompt plus optional scripts, packaged capability.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;RAG&lt;/strong&gt;: memory, agent stores custom data and pulls relevant chunks into
context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;AGENTS.md&lt;/strong&gt; / &lt;strong&gt;CLAUDE.md&lt;/strong&gt; / &lt;strong&gt;SOUL.md&lt;/strong&gt;: custom instructions loaded into
the context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context&lt;/strong&gt;: text the agent handles at once. User input, custom instructions,
tool outputs, RAG-retrieved chunks. Bounded by the model's context length.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Providers are wrapping yesterday's chats in agent loops. Execution pattern&lt;br&gt;
flips, the chat UI stays. Same split still holds: model caps capability, harness&lt;br&gt;
wires integrations and workflow, and intent has to be decomposed into pieces the&lt;br&gt;
agent can handle. Who does the decomposition is the next question.&lt;/p&gt;
&lt;h2&gt;
  
  
  AI platforms and architectures
&lt;/h2&gt;

&lt;p&gt;First, split common LLM application architectures by workflow: the execution&lt;br&gt;
pattern around model calls.&lt;/p&gt;

&lt;p&gt;At tiers 0-3 this is mostly application code. Fixed calls, branching, one-off&lt;br&gt;
tool use. At tiers 4-5 it becomes an agent harness. Loop, state, permissions,&lt;br&gt;
memory, orchestration.&lt;/p&gt;

&lt;p&gt;Six tiers, in order of escalating cost/complexity:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;Pattern&lt;/th&gt;
&lt;th&gt;When&lt;/th&gt;
&lt;th&gt;Cost shape&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;Single prompt&lt;/td&gt;
&lt;td&gt;Text in, text out&lt;/td&gt;
&lt;td&gt;1 LLM call&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Prompt chain&lt;/td&gt;
&lt;td&gt;Multi-step but predictable pipeline&lt;/td&gt;
&lt;td&gt;N LLM calls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Routing&lt;/td&gt;
&lt;td&gt;Input-type dispatch into one of K branches&lt;/td&gt;
&lt;td&gt;Router call + selected branch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Tool-use, single round&lt;/td&gt;
&lt;td&gt;Need external data, one tool round suffices&lt;/td&gt;
&lt;td&gt;Tool schema tokens + tool call + synthesis&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Agent, ReAct loop&lt;/td&gt;
&lt;td&gt;Iterative, unknown depth, state-dependent&lt;/td&gt;
&lt;td&gt;Unbounded: loop iterations × model/tool calls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Multi-agent / orchestrator&lt;/td&gt;
&lt;td&gt;Long-horizon, decomposable, large context&lt;/td&gt;
&lt;td&gt;Orchestrator + workers + replans + synthesis&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This post is about tiers 4 and 5. Agent loops, orchestrators, and who owns&lt;br&gt;
intent decomposition.&lt;/p&gt;

&lt;p&gt;Three families compete for "AI agent" today. Same lens for each: who owns intent&lt;br&gt;
decomposition? The specialized harness, the personal assistant, or me.&lt;/p&gt;
&lt;h3&gt;
  
  
  Specialized AI agents
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;AI trading agent&lt;/li&gt;
&lt;li&gt;AI legal assistant&lt;/li&gt;
&lt;li&gt;Customer service bot&lt;/li&gt;
&lt;li&gt;AI for medical diagnosis&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same product shape. Narrow task, fixed workflow, integrations hidden behind the&lt;br&gt;
UI.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Perplexity&lt;/em&gt; is the research/search version of that shape.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Research the latest evidence and debate around emergent agency in AI systems&lt;br&gt;
in 2025-2026.&lt;/p&gt;

&lt;blockquote&gt;
&lt;ul&gt;
&lt;li&gt;Searching the web

&lt;ul&gt;
&lt;li&gt;... some sources ...&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Searching recent academic and policy sources on emergent agency in AI
systems&lt;/li&gt;
&lt;li&gt;Retrieving recent reports and policy documents on AI agent behavior and
emergent agency&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;/blockquote&gt;

&lt;p&gt;Fine results, at least looks relevant.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;what's the model you using?&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;I'm Perplexity, powered by Grok 4.1. Perplexity's default model is a&lt;br&gt;
custom-optimized version called Sonar, built on Llama 3.1 or later&lt;br&gt;
iterations like Llama 3.3 70B, fine-tuned for fast web search,&lt;br&gt;
summarization, and efficiency. It's available to both free and Pro users for&lt;br&gt;
quick, relevant results. Pro subscribers can switch to advanced options like&lt;br&gt;
GPT-5.2 or Claude models via the model selector.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;/blockquote&gt;

&lt;p&gt;Fail. It understood the surface intent (answer a product-stack question), but&lt;br&gt;
the product owned routing and context. I had no way to inspect which model&lt;br&gt;
routed the answer, which metadata it saw, or why it mixed Grok, Sonar, and Llama&lt;br&gt;
into one pile.&lt;/p&gt;

&lt;p&gt;The outcome: specialized harness frames intent into its fixed shape. When that&lt;br&gt;
frame fits, I get a clean research answer. When the frame itself is wrong, I get&lt;br&gt;
confident product salad and no useful control surface.&lt;/p&gt;
&lt;h3&gt;
  
  
  Personal assistant
&lt;/h3&gt;

&lt;p&gt;Tried &lt;em&gt;OpenClaw&lt;/em&gt; first. It wants 2+ CPUs and 8+ GB RAM; my VPS has 1 and 1. Ran&lt;br&gt;
it anyway. It choked the VPS. Dropped it for &lt;em&gt;Hermes&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Rarely discussed, but experimental software with a lot of external integrations&lt;br&gt;
has too broad an attack surface, see &lt;a href="https://days-since-openclaw-cve.com" rel="noopener noreferrer"&gt;https://days-since-openclaw-cve.com&lt;/a&gt;. Keep&lt;br&gt;
it in mind.&lt;/p&gt;
&lt;h4&gt;
  
  
  Hermes, an agent that grows with you
&lt;/h4&gt;

&lt;p&gt;Strange that Nous Research doesn't mention they have a Docker image,&lt;br&gt;
&lt;code&gt;docker.io/nousresearch/hermes-agent&lt;/code&gt;, which I've successfully set up in&lt;br&gt;
&lt;a href="https://bogomolov.work/blog/posts/the-actual-state-of-self-hosting-on-a-vps/" rel="noopener noreferrer"&gt;podman&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  /etc/containers/systemd/hermes-agent.container
  &lt;br&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ini"&gt;&lt;code&gt;&lt;span class="nn"&gt;[Unit]&lt;/span&gt;
&lt;span class="py"&gt;Description&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;Hermes Agent&lt;/span&gt;
&lt;span class="py"&gt;Wants&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;network-online.target&lt;/span&gt;
&lt;span class="py"&gt;After&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;network-online.target&lt;/span&gt;

&lt;span class="nn"&gt;[Container]&lt;/span&gt;
&lt;span class="py"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;docker.io/nousresearch/hermes-agent:latest&lt;/span&gt;
&lt;span class="py"&gt;ContainerName&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;hermes-agent&lt;/span&gt;
&lt;span class="py"&gt;Network&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;selfhosted&lt;/span&gt;
&lt;span class="py"&gt;Volume&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;/root/hermes:/opt/data&lt;/span&gt;
&lt;span class="py"&gt;Volume&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;/root/hermes-root:/root&lt;/span&gt;
&lt;span class="py"&gt;Volume&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;/tmp/hermes:/tmp&lt;/span&gt;
&lt;span class="py"&gt;Ulimit&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;nofile=1024:1024&lt;/span&gt;

&lt;span class="py"&gt;Environment&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;VIRTUAL_ENV=/root/.venv&lt;/span&gt;
&lt;span class="py"&gt;Environment&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;PYTHONPATH=/root/.venv/lib/python3.13/site-packages&lt;/span&gt;

&lt;span class="py"&gt;Exec&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;gateway run&lt;/span&gt;

&lt;span class="nn"&gt;[Service]&lt;/span&gt;
&lt;span class="py"&gt;Restart&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;always&lt;/span&gt;
&lt;span class="py"&gt;RestartSec&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;3&lt;/span&gt;
&lt;span class="py"&gt;MemoryMax&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;768M&lt;/span&gt;
&lt;span class="py"&gt;MemorySwapMax&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;768M&lt;/span&gt;
&lt;span class="py"&gt;CPUQuota&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;85%&lt;/span&gt;
&lt;span class="py"&gt;TasksMax&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;128&lt;/span&gt;

&lt;span class="nn"&gt;[Install]&lt;/span&gt;
&lt;span class="py"&gt;WantedBy&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="s"&gt;multi-user.target&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;




&lt;p&gt;&lt;/p&gt;

&lt;p&gt;And it runs completely fine on a 1 CPU / 1 GB VPS.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  CPU/RAM consumption
  &lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft6u8bh7roho2ys916e9o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft6u8bh7roho2ys916e9o.png" alt="CPU/RAM consumption grafana screenshot" width="800" height="425"&gt;&lt;/a&gt;



&lt;p&gt;&lt;/p&gt;

&lt;p&gt;I connected it to my GPT subscription, added integrations for X, Google&lt;br&gt;
Calendar, Notion, and this blog's RSS, plus free Mem0 as RAG.&lt;/p&gt;

&lt;p&gt;It even worked right after setup, but the next day it forgot about the&lt;br&gt;
integration. I had to persuade it to try again.&lt;/p&gt;

&lt;p&gt;OAuth failed in a different way. During Google Calendar integration I issued&lt;br&gt;
credentials only for read/write on the calendar, not broader Google scopes. The&lt;br&gt;
builtin Google skill wants broader access, so the agent re-requests broader&lt;br&gt;
scopes every time it touches the calendar, and eventually the auth flow breaks&lt;br&gt;
again.&lt;/p&gt;

&lt;p&gt;One more case: I configured a scheduled job to check, each morning at 9:00, my&lt;br&gt;
Notion calendar, Google Calendar, event listings, and send me a summary for&lt;br&gt;
today and tomorrow. How often does it work right? Almost never. It checks only&lt;br&gt;
one calendar, sends events for the next ~6 months instead of 2 days, sends&lt;br&gt;
events for the current month but from this and previous years, and so on.&lt;/p&gt;

&lt;p&gt;Current state: the initial GPT auth token has expired, and the agent can't renew&lt;br&gt;
it automatically. Well... experiment successful.&lt;/p&gt;

&lt;p&gt;Each failure is easy to fix manually. Cron, small script, explicit OAuth scopes,&lt;br&gt;
date windows, deterministic calendar queries. But that is exactly the point: the&lt;br&gt;
general assistant is supposed to replace the glue. Here it doesn't. The failure&lt;br&gt;
is not the model. The harness decomposes intent badly, and the UI doesn't expose&lt;br&gt;
decomposition early enough to fix it.&lt;sup id="fnref3"&gt;3&lt;/sup&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  CLI coding agents
&lt;/h3&gt;

&lt;p&gt;Terminal-native, actively evolving. Everything in my hands. Only vendor ToS can&lt;br&gt;
limit me.&lt;/p&gt;
&lt;h4&gt;
  
  
  OpenCode, the open source AI coding agent
&lt;/h4&gt;

&lt;p&gt;My favorite one. Open source, standard &lt;code&gt;~/.config/opencode&lt;/code&gt; path, strong&lt;br&gt;
build/plan sub-agent architecture. Also ships &lt;code&gt;opencode web&lt;/code&gt;, same engine,&lt;br&gt;
browser UI.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  ~/.config/opencode/opencode.jsonc
  &lt;br&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"$schema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://opencode.ai/config.json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"autoupdate"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"default_agent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"plan"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"share"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"disabled"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"snapshot"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"instructions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"/Users/ivan/.config/opencode/AGENTS.md"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mcp"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"context7"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"remote"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://mcp.context7.com/mcp"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"headers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"CONTEXT7_API_KEY"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"{env:CONTEXT7_API_KEY}"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"enabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"playwright"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"local"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"npx"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"-y"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"@playwright/mcp@latest"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"--browser=chromium"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"--executable-path=/Applications/Chromium.app/Contents/MacOS/Chromium"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="s2"&gt;"--caps=vision,devtools"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"environment"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"PLAYWRIGHT_BROWSERS_PATH"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"{env:HOME}/.cache/ms-playwright"&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"enabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"experimental"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"openTelemetry"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"server"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"mdns"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"plugin"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;




&lt;p&gt;&lt;/p&gt;

&lt;p&gt;Current AI-bro consensus: "Context is the key". I agree. Context engineering&lt;br&gt;
pwnd &lt;a href="https://bogomolov.work/blog/posts/prompt-engineering-notes/" rel="noopener noreferrer"&gt;prompt engineering&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Even more, after&lt;br&gt;
&lt;a href="https://arxiv.org/abs/2602.11988" rel="noopener noreferrer"&gt;Gloaguen et al., "Evaluating AGENTS.md" (Feb 2026)&lt;/a&gt;,&lt;br&gt;
generated agent context files hurt task success and add ~20% inference cost. I&lt;br&gt;
stopped keeping per-repo CLAUDE.md / AGENTS.md files full of paths, framework&lt;br&gt;
summaries, and obvious project descriptions. Current agents can inspect a repo&lt;br&gt;
per case, fast and cheap enough. Put only what they cannot infer from code.&lt;br&gt;
Constraints, preferences, dangerous commands, external contracts.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  my global CLAUDE.md / AGENTS.md
  &lt;br&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Invariants&lt;/span&gt;

Be brief

Truth over comfort

Simple over clever

Contradiction: name both sides, never average

Use subagents to do the work; main thread is orchestrator

grep -&amp;gt; rg; python -&amp;gt; python3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;




&lt;p&gt;&lt;/p&gt;

&lt;p&gt;Still, the most reliable way to fix hallucinations is not to argue with the&lt;br&gt;
agent but to drop the session and restart from scratch.&lt;/p&gt;

&lt;p&gt;OpenCode's approach helps me clearly follow the principles above:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;In plan mode I explain the goal and constraints, pointing to important files
or docs as entry points;&lt;/li&gt;
&lt;li&gt;The plan sub-agent collects requirements, explores the repo, clarifies
intent, produces a plan&lt;sup id="fnref4"&gt;4&lt;/sup&gt;;&lt;/li&gt;
&lt;li&gt;I review the plan; if decomposition is wrong, GOTO 1;&lt;/li&gt;
&lt;li&gt;The approved plan is handed to the build sub-agent, which implements it&lt;sup id="fnref5"&gt;5&lt;/sup&gt;;&lt;/li&gt;
&lt;li&gt;GOTO 1.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In plan mode, the harness exposes decomposition before execution. At that point,&lt;br&gt;
the limit is me: how clearly I can split intent into steps.&lt;/p&gt;

&lt;p&gt;This resembles the Plan-Then-Execute pattern&lt;sup id="fnref6"&gt;6&lt;/sup&gt;.&lt;/p&gt;
&lt;h4&gt;
  
  
  Claude Code
&lt;/h4&gt;

&lt;p&gt;I migrated Claude work to Claude Code after Anthropic restricted third-party&lt;br&gt;
Claude access to API credits&lt;sup id="fnref7"&gt;7&lt;/sup&gt;; OpenCode still runs my GPT subscription.&lt;/p&gt;

&lt;p&gt;Proprietary, vendor-locked, but I can't complain that it misses anything&lt;br&gt;
important.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  ~/.claude/settings.json
  &lt;br&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"$schema"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://json.schemastore.org/claude-code-settings.json"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cleanupPeriodDays"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"env"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"CLAUDE_CODE_DISABLE_FEEDBACK_SURVEY"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"CLAUDE_CODE_DISABLE_TERMINAL_TITLE"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"DISABLE_AUTOUPDATER"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"DISABLE_BUG_COMMAND"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"DISABLE_COST_WARNINGS"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"DISABLE_ERROR_REPORTING"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"DISABLE_FEEDBACK_COMMAND"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"DISABLE_TELEMETRY"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"1"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"permissions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"allow"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git log *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git diff *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git show *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git status)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(grep *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(echo *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(ls *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(rg *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(npm run *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(npm test *)"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"deny"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(curl *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(docker push *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(find * -delete)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(find * -exec rm*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git branch -D *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git checkout -- *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git clean -f*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git push *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(git reset --hard*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(nc *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(rm -f *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(rm -r *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(rm -rf *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(rsync *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(scp *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(ssh *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(sudo *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Bash(wget *)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Edit(./.env*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Edit(./.git/**)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Edit(./secrets/**)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Edit(~/.aws/**)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Edit(~/.bashrc)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Edit(~/.ssh/**)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Edit(~/.zshrc)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(*.env)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(./.env.*)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(./secrets/**)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(~/.aws/**)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(~/.azure/**)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(~/.config/gh/**)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(~/.git-credentials)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(~/.gnupg/**)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(~/.kube/**)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(~/.npmrc)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"Read(~/.ssh/**)"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"defaultMode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"plan"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sonnet"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"disableClaudeAiConnectors"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"disableBundledSkills"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"disableRemoteControl"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"disableWorkflows"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"disableArtifact"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"enableAllProjectMcpServers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"includeCoAuthoredBy"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"sandbox"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"enabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"excludedCommands"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"git"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"effortLevel"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"low"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"awaySummaryEnabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"autoUpdatesChannel"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"stable"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"autoMemoryEnabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"disableAutoMode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"disable"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"theme"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"light"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"editorMode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"normal"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"preferredNotifChannel"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"notifications_disabled"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"autoCompactEnabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"skipAutoPermissionPrompt"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;




&lt;p&gt;&lt;/p&gt;

&lt;p&gt;This config reaches the same control goal I like in OpenCode: decomposition&lt;br&gt;
stays visible, writes stay gated, mode switching stays under my control:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;"defaultMode": "plan"&lt;/code&gt;: plan mode by default. Writes blocked, reads allowed,
plan exposed as &lt;code&gt;/plan&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;"disableAutoMode": "disable"&lt;/code&gt;: no autonomous mode switching&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;"sandbox.enabled": true&lt;/code&gt; plus a &lt;code&gt;permissions.deny&lt;/code&gt; list for &lt;code&gt;rm -rf&lt;/code&gt;,
&lt;code&gt;git push&lt;/code&gt;, &lt;code&gt;sudo&lt;/code&gt;, &lt;code&gt;~/.ssh/**&lt;/code&gt;, etc.&lt;/li&gt;
&lt;li&gt;telemetry off: &lt;code&gt;DISABLE_TELEMETRY=1&lt;/code&gt;, &lt;code&gt;DISABLE_FEEDBACK_COMMAND=1&lt;/code&gt;,
&lt;code&gt;DISABLE_ERROR_REPORTING=1&lt;/code&gt;, &lt;code&gt;awaySummaryEnabled: false&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h5&gt;
  
  
  Current setup
&lt;/h5&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Context7 MCP&lt;/strong&gt;: provides actual library docs instead of the model's stale or
hallucinated snippets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Playwright MCP&lt;/strong&gt;: browser control. Nothing else to add.&lt;sup id="fnref8"&gt;8&lt;/sup&gt;
&lt;/li&gt;
&lt;li&gt;~&lt;strong&gt;Caveman plugin&lt;/strong&gt;: its selling point is "Saves tokens, preserve
accuracy".~&lt;sup id="fnref9"&gt;9&lt;/sup&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;security-guidance@claude-plugins-official&lt;/strong&gt;: blame the times&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each of these touches the harness only. Model stays the vendor's, intent stays&lt;br&gt;
mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Intent is the failure point
&lt;/h2&gt;

&lt;p&gt;Perplexity frames intent into its fixed shape. Hermes decomposes intent on its&lt;br&gt;
own, without me. CLI plan mode keeps decomposition visible. I'm the bottleneck.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The loop is what makes something agentic&lt;/em&gt;, and the harness puts me inside or&lt;br&gt;
outside of it.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Outside the loop: I wait for a solution that fits my intent's shape. When it
doesn't, the agent won't say "no". It attempts the task anyway, drifts, and
hands something back.&lt;/li&gt;
&lt;li&gt;Inside the loop: I keep intent, decompose it and approve execution. The
ceiling is whatever I can break into steps.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Same year. Same frontier models. Humans still hold the loop. That is the&lt;br&gt;
autonomy that emerges today.&lt;sup id="fnref10"&gt;10&lt;/sup&gt;&lt;/p&gt;




&lt;ol&gt;

&lt;li id="fn1"&gt;
&lt;p&gt;Popularized by ReAct: Yao et al., ICLR 2023&lt;br&gt;
(&lt;a href="https://arxiv.org/abs/2210.03629" rel="noopener noreferrer"&gt;paper&lt;/a&gt;,&lt;br&gt;
&lt;a href="https://research.google/blog/react-synergizing-reasoning-and-acting-in-language-models/" rel="noopener noreferrer"&gt;Google Research blog&lt;/a&gt;).&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn2"&gt;
&lt;p&gt;See Birgitta Böckeler, "Harness engineering for coding agent users":&lt;br&gt;
&lt;a href="https://martinfowler.com/articles/harness-engineering.html" rel="noopener noreferrer"&gt;https://martinfowler.com/articles/harness-engineering.html&lt;/a&gt;.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn3"&gt;
&lt;p&gt;&lt;em&gt;Personal assistants&lt;/em&gt; still look worth watching. I'll wait for the next&lt;br&gt;
Hermes iteration. Notion's agent may become the boring alternative.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn4"&gt;
&lt;p&gt;Yes-yes, I know about openspec.dev, but the plan has to stay observable, not&lt;br&gt;
5+ A4 &lt;a href="https://bogomolov.work/blog/posts/rotten-specs/" rel="noopener noreferrer"&gt;neuro-generated&lt;/a&gt; pages of&lt;br&gt;
raw text.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn5"&gt;
&lt;p&gt;To save some tokens I use stronger model for plan (Opus) and weaker for&lt;br&gt;
build (Sonnet).&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn6"&gt;
&lt;p&gt;&lt;a href="https://simonwillison.net/2025/Jun/13/prompt-injection-design-patterns/#the-plan-then-execute-pattern" rel="noopener noreferrer"&gt;https://simonwillison.net/2025/Jun/13/prompt-injection-design-patterns/#the-plan-then-execute-pattern&lt;/a&gt;&lt;br&gt;
Caveat: the split is weaker here. Plan sub-agent still reads untrusted repo&lt;br&gt;
while planning, so a malicious file can steer the plan.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn7"&gt;
&lt;p&gt;&lt;a href="https://bogomolov.work/blog/posts/will-ai-replace-developers/" rel="noopener noreferrer"&gt;Anthropic&lt;/a&gt;&lt;br&gt;
tightened terms. Subscription plans (including the corporate one I use) no&lt;br&gt;
longer cover Claude access from third-party apps; third-party access now&lt;br&gt;
requires per-usage API credits.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn8"&gt;
&lt;p&gt;CLI is more token effective than MCP,&lt;br&gt;
&lt;a href="https://github.com/microsoft/playwright-cli#playwright-cli-vs-playwright-mcp" rel="noopener noreferrer"&gt;https://github.com/microsoft/playwright-cli#playwright-cli-vs-playwright-mcp&lt;/a&gt;.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn9"&gt;
&lt;p&gt;Caveman vs "be brief",&lt;br&gt;
&lt;a href="https://www.maxtaylor.me/articles/i-benchmarked-caveman-against-two-words" rel="noopener noreferrer"&gt;https://www.maxtaylor.me/articles/i-benchmarked-caveman-against-two-words&lt;/a&gt;.&lt;br&gt;
Agent tools move too fast; without a fresh benchmark after model or harness&lt;br&gt;
updates, a third-party skill can quietly make results worse than the&lt;br&gt;
baseline. Benchmark it continuously or keep it off.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn10"&gt;
&lt;p&gt;&lt;a href="https://www.anthropic.com/news/measuring-agent-autonomy" rel="noopener noreferrer"&gt;Agent autonomy&lt;/a&gt;&lt;br&gt;
framed as emergent from model behavior, product design, and user oversight&lt;br&gt;
strategy.&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>architecture</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
