<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: ProxySoul</title>
    <description>The latest articles on DEV Community by ProxySoul (@proxyosul).</description>
    <link>https://dev.to/proxyosul</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4126852%2Fda663301-4047-49d1-88ad-a3ad1016bdb6.JPG</url>
      <title>DEV Community: ProxySoul</title>
      <link>https://dev.to/proxyosul</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/proxyosul"/>
    <language>en</language>
    <item>
      <title>Empryo vs pi on Opus 5: four bugs fixed, $2.27 vs $3.10</title>
      <dc:creator>ProxySoul</dc:creator>
      <pubDate>Sun, 27 Sep 2026 16:55:13 +0000</pubDate>
      <link>https://dev.to/proxyosul/empryo-vs-pi-on-opus-5-four-bugs-fixed-227-vs-310-3eml</link>
      <guid>https://dev.to/proxyosul/empryo-vs-pi-on-opus-5-four-bugs-fixed-227-vs-310-3eml</guid>
      <description>&lt;p&gt;I build Empryo, an AI coding agent for desktop and terminal. In my benchmark published on 16 August 2026, Empryo and pi both fixed all four bugs on Opus 5. The published round totals were $2.27 for Empryo and $3.10 for pi, summing the per-task costs.&lt;/p&gt;

&lt;p&gt;I'm focusing on Opus because both agents passed every task in that slice, so the cost comparison doesn't trade away a successful fix. This isn't a claim about today's versions or every model. The wider report covers six models, and explicitly notes that pi can still be cheaper on inexpensive models.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the money went
&lt;/h2&gt;

&lt;p&gt;The bugs came from hono and opencode. Hidden regression tests from the merged fixes graded the changes.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Opus 5 task&lt;/th&gt;
&lt;th&gt;Empryo&lt;/th&gt;
&lt;th&gt;pi&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cookie mangling, hono&lt;/td&gt;
&lt;td&gt;$0.468&lt;/td&gt;
&lt;td&gt;$0.269&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;TrieRouter regexp, hono&lt;/td&gt;
&lt;td&gt;$0.570&lt;/td&gt;
&lt;td&gt;$0.634&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Symlinked grep root, opencode&lt;/td&gt;
&lt;td&gt;$0.762&lt;/td&gt;
&lt;td&gt;$1.19&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Revert by ID, opencode&lt;/td&gt;
&lt;td&gt;$0.472&lt;/td&gt;
&lt;td&gt;$1.01&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Round total, rounded&lt;/td&gt;
&lt;td&gt;$2.27&lt;/td&gt;
&lt;td&gt;$3.10&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;pi was cheaper on the cookie bug. Most of Empryo's net savings came from the two opencode tasks.&lt;/p&gt;

&lt;p&gt;The published report uses a small number of attempts: mostly two per result, with some single runs. Both repositories are TypeScript. Empryo ran with 37 tools and its Genome code map; the tested pi configuration used four tools. This compares those configurations, not every possible pi setup.&lt;/p&gt;

&lt;p&gt;The timing protocol also prebuilt Empryo's code index. Cold indexing was reported separately at 96-100 seconds for opencode and eight seconds for hono. Don't mistake an agent-working timer for a fresh-install timer. I'm keeping this post's comparison to cost and correctness.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://empryo.com/benchmarks/forge-v2/" rel="noopener noreferrer"&gt;Results, charts and limitations&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/proxysoul/empryo-bench/tree/main/round-3" rel="noopener noreferrer"&gt;Public tasks, raw results and reproduction setup&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>coding</category>
      <category>llm</category>
    </item>
    <item>
      <title>In Empryo, a bug fix should leave a guard behind</title>
      <dc:creator>ProxySoul</dc:creator>
      <pubDate>Sun, 27 Sep 2026 13:28:27 +0000</pubDate>
      <link>https://dev.to/proxyosul/in-empryo-a-bug-fix-should-leave-a-guard-behind-4ol5</link>
      <guid>https://dev.to/proxyosul/in-empryo-a-bug-fix-should-leave-a-guard-behind-4ol5</guid>
      <description>&lt;p&gt;i don't want to fix the same bug again next week because another agent wrote the same pattern. so this is the rule in Empryo, an AI coding agent i'm building: a fix should leave a guard behind!&lt;/p&gt;

&lt;p&gt;i call this the immune system. here's the useful part if u want to build one for ur own project.&lt;/p&gt;

&lt;h2&gt;
  
  
  reproduce before fixing
&lt;/h2&gt;

&lt;p&gt;a hunter gets a slice of the code and one kind of bug to look for. it has to reproduce the failure in a throwaway home directory, away from my real config. a convincing paragraph about a possible bug isn't enough.&lt;/p&gt;

&lt;p&gt;an independent reviewer reruns that reproduction. if it doesn't hold up, the report gets rejected. the agent that found it doesn't get to approve its own finding.&lt;/p&gt;

&lt;h2&gt;
  
  
  prove the fix, then guard the pattern
&lt;/h2&gt;

&lt;p&gt;the fixer has to show the failure on the old code and the passing case on the new code. another reviewer checks the diff and reruns the proof.&lt;/p&gt;

&lt;p&gt;for patterns i can catch statically, the fix also gets a GritQL rule in Biome. two fixtures go with it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;bad code the rule must flag&lt;/li&gt;
&lt;li&gt;good code the rule must leave alone&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;then i run the rule over the production code. if the pattern exists somewhere else, there's more work to do.&lt;/p&gt;

&lt;p&gt;that's what i want from a fix. the next agent shouldn't need to remember a warning buried in an old conversation. lint should catch the pattern when it writes it again.&lt;/p&gt;

&lt;p&gt;a lint rule still has limits! it catches the shape i taught it, not every possible version of the bug. the reproduction and real-app checks still matter.&lt;/p&gt;

&lt;h2&gt;
  
  
  keep the handoffs in files
&lt;/h2&gt;

&lt;p&gt;a bug is a markdown file. its folder is its state: found, ready, fixed, rejected. an agent writes its evidence there before handing it off.&lt;/p&gt;

&lt;p&gt;claiming a record uses &lt;code&gt;mkdir claims/&amp;lt;id&amp;gt;&lt;/code&gt;. only the agent whose mkdir succeeds owns that claim. stale claims still need handling when a worker dies.&lt;/p&gt;

&lt;p&gt;the record survives the session. and anything disputed comes back to human triage. i'm not letting a pile of confident reports decide what's true.&lt;/p&gt;

&lt;h2&gt;
  
  
  take the skill
&lt;/h2&gt;

&lt;p&gt;the workflow is public in &lt;a href="https://github.com/proxysoul/SoulStack/tree/main/skills/immune-system" rel="noopener noreferrer"&gt;SoulStack's immune-system skill&lt;/a&gt;. u can use the approach with an agent that runs shell commands; it doesn't require switching coding agents.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://empryo.com" rel="noopener noreferrer"&gt;Empryo&lt;/a&gt; is the app i'm building it around. its Genome maps symbols, callers and imports, which gives the hunters somewhere concrete to start.&lt;/p&gt;

&lt;p&gt;start small: one reproduced bug, one reviewed fix, one guard u can prove. scale after that.&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
    <item>
      <title>How i give my coding agent a map of the repo with Empryo</title>
      <dc:creator>ProxySoul</dc:creator>
      <pubDate>Tue, 15 Sep 2026 19:54:13 +0000</pubDate>
      <link>https://dev.to/proxyosul/how-i-give-my-coding-agent-a-map-of-the-repo-with-empryo-2a3g</link>
      <guid>https://dev.to/proxyosul/how-i-give-my-coding-agent-a-map-of-the-repo-with-empryo-2a3g</guid>
      <description>&lt;p&gt;i built Empryo solo because i was tired of explaining the same repo to my coding agent. where things live, what calls what, which files usually need to change together. then a new session starts and i'm doing it again.&lt;/p&gt;

&lt;p&gt;Token usage bothered me too. A few prompts can involve a lot of reading, retries, and tool results. The prompt you type is only a small part of what the model processes.&lt;/p&gt;

&lt;p&gt;So i started with the thing i wanted the agent to have before it opened another file: a useful map.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Genome
&lt;/h2&gt;

&lt;p&gt;The screenshot above is Empryo's Genome dock. It maps files, symbols, and imports, and updates as the code changes. The little agent moves around the graph while it works. yes, i gave it a home.&lt;/p&gt;

&lt;p&gt;Behind that view, the agent gets a ranked text map. Connections between files matter. Recent reads and edits matter. Files that tend to change together in Git can provide another clue.&lt;/p&gt;

&lt;p&gt;The graph helps choose where to look. The agent still needs to open the source, check its assumptions, and run the relevant tests.&lt;/p&gt;

&lt;h2&gt;
  
  
  One rename, several places to check
&lt;/h2&gt;

&lt;p&gt;Take an illustrative task: rename a session method and update its callers.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;auth/session.ts
  refreshSession()
       |
       +-- api/client.ts
       +-- workers/refresh.ts
       +-- tests/session.test.ts
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A filename search might find the implementation immediately. The more interesting question is what depends on it.&lt;/p&gt;

&lt;p&gt;An import relationship gives the agent a place to investigate. A test that often changes alongside the implementation is another lead. Neither proves that every caller has been found. Dynamic calls, generated code, and string references still need checking.&lt;/p&gt;

&lt;p&gt;The workflow i want is simple:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Use the map to pick the likely implementation and dependents.&lt;/li&gt;
&lt;li&gt;Read those files before changing the interface.&lt;/li&gt;
&lt;li&gt;Make the change and update the map.&lt;/li&gt;
&lt;li&gt;Search for remaining references and run the focused tests.&lt;/li&gt;
&lt;li&gt;Review the diff for changes outside the task.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Keeping that map current matters. A map of yesterday's code can confidently send an agent to the wrong place.&lt;/p&gt;

&lt;h2&gt;
  
  
  Same idea, in the terminal
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgkpwvi3alanh4h9ce8n8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgkpwvi3alanh4h9ce8n8.png" alt="Empryo terminal UI showing changed files, dependents, Git co-change hints, and the pink Mote mascot" width="800" height="520"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the TUI. The change view shows dependents and co-change hints beside the files. i want those clues visible to me too, so i can see what the agent is working from.&lt;/p&gt;

&lt;p&gt;It also has themes. i spend enough time in a terminal to care how it feels to sit in one.&lt;/p&gt;

&lt;h2&gt;
  
  
  A cheaper model can still make a more expensive run
&lt;/h2&gt;

&lt;p&gt;Empryo lets me route different tasks to different models. Exploration, editing, and review don't always need the same choice.&lt;/p&gt;

&lt;p&gt;But switching models has a catch: cache reuse.&lt;/p&gt;

&lt;p&gt;A worker on the same model can sometimes reuse an existing prompt prefix. Moving it to another model can mean paying to process that context again. A lower token price doesn't automatically mean a lower total.&lt;/p&gt;

&lt;p&gt;This is why i care about cached input, fresh input, output, retries, and whether the change actually worked. A failed cheap attempt followed by a full retry is still two attempts.&lt;/p&gt;

&lt;p&gt;Subscription usage is another separate measure. A token-price estimate is not the amount charged to a subscription, and cache percentage is not a percentage of money saved.&lt;/p&gt;

&lt;h2&gt;
  
  
  A small test i'd actually trust
&lt;/h2&gt;

&lt;p&gt;If you're comparing coding agents, give each one the same commit and a small task with a checkable result. Keep the model and permissions comparable where possible.&lt;/p&gt;

&lt;p&gt;Record:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Check&lt;/th&gt;
&lt;th&gt;What to write down&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Correctness&lt;/td&gt;
&lt;td&gt;Did the relevant tests pass, and does the diff solve the task?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Exploration&lt;/td&gt;
&lt;td&gt;Which files were read before the first edit?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rework&lt;/td&gt;
&lt;td&gt;How many retries or corrections were needed?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Usage&lt;/td&gt;
&lt;td&gt;Fresh input, cached input, output, and the billing basis&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Your time&lt;/td&gt;
&lt;td&gt;How much explanation and review did it take?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Repeat it on a few tasks. One attractive screenshot can't tell you which setup works best for your repo.&lt;/p&gt;

&lt;p&gt;That's the problem i'm working on with &lt;a href="https://empryo.com" rel="noopener noreferrer"&gt;Empryo&lt;/a&gt;: better project awareness, with desktop and terminal interfaces that make the work easier to follow.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/proxysoul/empryo" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;How do you handle this today? A repo instructions file, your own index, or a lot of repeated explanations? i'm especially interested in what breaks once a project gets big.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;AI disclosure: this article was drafted by an AI agent using my project documentation, screenshots, and the notes i supplied.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
  </channel>
</rss>
