<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Harjot Rana</title>
    <description>The latest articles on DEV Community by Harjot Rana (@harjjotsinghh).</description>
    <link>https://dev.to/harjjotsinghh</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1193425%2Fb5f7b30b-df4d-4bac-9732-6b1e1ac7805a.jpg</url>
      <title>DEV Community: Harjot Rana</title>
      <link>https://dev.to/harjjotsinghh</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/harjjotsinghh"/>
    <language>en</language>
    <item>
      <title>How to run Muse Code on a remote server over SSH, with a desktop GUI</title>
      <dc:creator>Harjot Rana</dc:creator>
      <pubDate>Mon, 28 Sep 2026 04:59:12 +0000</pubDate>
      <link>https://dev.to/harjjotsinghh/how-to-run-muse-code-on-a-remote-server-over-ssh-with-a-desktop-gui-3h1o</link>
      <guid>https://dev.to/harjjotsinghh/how-to-run-muse-code-on-a-remote-server-over-ssh-with-a-desktop-gui-3h1o</guid>
      <description>&lt;p&gt;Your code lives on a server: a beefy dev box, a cloud VM, the machine under your desk. You want Muse Code to work there, where the files, the toolchain and the CPU are. But you'd rather not spend the day in an SSH terminal scrolling back through agent output.&lt;/p&gt;

&lt;p&gt;Helicon 0.18 adds &lt;strong&gt;SSH projects&lt;/strong&gt;: you pick a folder on a remote machine, and Helicon runs Muse Code there over &lt;code&gt;ssh&lt;/code&gt;, while you get the desktop app on your laptop: threads in a sidebar, readable approvals, and inline diffs.&lt;/p&gt;

&lt;p&gt;This guide sets it up in about 10 minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it works
&lt;/h2&gt;

&lt;p&gt;When you start a thread in an SSH project, Helicon runs this on your laptop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ssh &amp;lt;host&amp;gt; &lt;span class="s1"&gt;'cd /path/to/project &amp;amp;&amp;amp; exec muse serve'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Muse Code runs on the server, inside the project folder, and talks to Helicon over the SSH connection using the Muse Session Protocol (MSP), the same protocol it speaks to any local client. Nothing is synced or copied: the code never leaves the server.&lt;/p&gt;

&lt;p&gt;A few consequences worth knowing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The server's Muse login is the one used.&lt;/strong&gt; Whatever account you signed into with &lt;code&gt;muse login&lt;/code&gt; on the server is what your SSH threads run on. Helicon's local account profiles don't apply to SSH projects.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your laptop needs to stay connected.&lt;/strong&gt; Muse runs as part of the SSH session, so if your laptop sleeps or the network drops, that session ends. Helicon notices within about 45 seconds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your SSH config is respected.&lt;/strong&gt; Host aliases, ports, users, jump hosts and keys all come from &lt;code&gt;~/.ssh/config&lt;/code&gt;, the same as when you type &lt;code&gt;ssh&lt;/code&gt; yourself.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  1. Make &lt;code&gt;ssh&lt;/code&gt; work without a password
&lt;/h2&gt;

&lt;p&gt;Helicon never types a password or answers a prompt; it runs &lt;code&gt;ssh&lt;/code&gt; in batch mode. So the first step is a key.&lt;/p&gt;

&lt;p&gt;On your laptop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ssh-keygen &lt;span class="nt"&gt;-t&lt;/span&gt; ed25519            &lt;span class="c"&gt;# skip if you already have ~/.ssh/id_ed25519&lt;/span&gt;
ssh-copy-id you@devbox.example.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then give the host a short name in &lt;code&gt;~/.ssh/config&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ssh"&gt;&lt;code&gt;&lt;span class="k"&gt;Host&lt;/span&gt; devbox
  &lt;span class="k"&gt;HostName&lt;/span&gt; devbox.example.com
  &lt;span class="k"&gt;User&lt;/span&gt; you
  &lt;span class="k"&gt;IdentityFile&lt;/span&gt; ~/.ssh/id_ed25519
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Check it: this should print &lt;code&gt;ok&lt;/code&gt; without asking you anything.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ssh devbox &lt;span class="nb"&gt;echo &lt;/span&gt;ok
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If it asks you to confirm the host key, answer &lt;code&gt;yes&lt;/code&gt; once. Helicon can't answer that question for you, and will tell you so if you skip it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;On Windows:&lt;/strong&gt; Helicon uses Windows' own OpenSSH (&lt;code&gt;ssh.exe&lt;/code&gt;), not the one inside WSL. Keys and config belong in &lt;code&gt;%USERPROFILE%\.ssh&lt;/code&gt;. If your keys only live in WSL, copy them over, or run &lt;code&gt;ssh-keygen&lt;/code&gt; in PowerShell.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Install Muse Code on the server
&lt;/h2&gt;

&lt;p&gt;On the server, install the &lt;code&gt;muse&lt;/code&gt; CLI the way you normally would, then sign in once:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;muse login
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now make sure &lt;code&gt;muse&lt;/code&gt; is on the PATH for &lt;strong&gt;non-interactive&lt;/strong&gt; SSH commands, which is what Helicon uses. From your laptop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ssh devbox &lt;span class="s1"&gt;'command -v muse'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If that prints nothing, your shell only sets PATH for interactive logins. Add the folder that holds &lt;code&gt;muse&lt;/code&gt; to your PATH in &lt;code&gt;~/.zshenv&lt;/code&gt; (zsh) or near the top of &lt;code&gt;~/.bashrc&lt;/code&gt; (bash), and try again.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Add the project in Helicon
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Open Helicon (0.18 or newer) and click &lt;strong&gt;Add project&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Choose &lt;strong&gt;SSH host&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Type the host: &lt;code&gt;devbox&lt;/code&gt;, or &lt;code&gt;you@devbox.example.com&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Browse the server's folders, starting at your home folder, and pick the project.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The project shows up in the sidebar like any other. Start a thread, and Muse Code is working on the server.&lt;/p&gt;

&lt;h2&gt;
  
  
  When something goes wrong
&lt;/h2&gt;

&lt;p&gt;Helicon reports what &lt;code&gt;ssh&lt;/code&gt; itself said, so the error usually tells you the fix:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Helicon says&lt;/th&gt;
&lt;th&gt;Fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Host key verification failed&lt;/td&gt;
&lt;td&gt;Run &lt;code&gt;ssh devbox&lt;/code&gt; once in a terminal and accept the host key&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Permission denied (publickey)&lt;/td&gt;
&lt;td&gt;The key isn't on the server: rerun &lt;code&gt;ssh-copy-id&lt;/code&gt;, or check &lt;code&gt;IdentityFile&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Could not find ssh on this machine&lt;/td&gt;
&lt;td&gt;Install OpenSSH (on Windows: Settings, Optional features, OpenSSH Client)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Connection refused / timed out&lt;/td&gt;
&lt;td&gt;Check the host, port and VPN with &lt;code&gt;ssh devbox echo ok&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Muse not found when starting a thread&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;ssh devbox 'command -v muse'&lt;/code&gt; prints nothing: fix PATH as in step 2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What SSH projects don't do yet
&lt;/h2&gt;

&lt;p&gt;In 0.18, SSH projects run threads, show diffs and approvals, and list past sessions from the server. These aren't available for them yet:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the file viewer panel&lt;/li&gt;
&lt;li&gt;the skills list&lt;/li&gt;
&lt;li&gt;attaching files (images still work)&lt;/li&gt;
&lt;li&gt;creating a new remote folder or cloning a repo onto the server from Helicon&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Create folders and clones over SSH first, then add them.&lt;/p&gt;

&lt;h2&gt;
  
  
  SSH projects or the remote daemon?
&lt;/h2&gt;

&lt;p&gt;Helicon has two ways to use a remote machine:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SSH projects&lt;/strong&gt; (this guide): the Helicon app runs on your laptop, and only Muse runs on the server. Mix local and remote projects in one window.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;a href="https://helicon.sh/features/remote-daemon" rel="noopener noreferrer"&gt;Remote daemon&lt;/a&gt;&lt;/strong&gt;: the whole Helicon server runs on the remote machine, and you open it in a browser. Better when you want to reach it from any device.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;Helicon is free and open source (MIT) for Windows, macOS and Linux: &lt;a href="https://helicon.sh/?utm_source=devto&amp;amp;utm_medium=article&amp;amp;utm_campaign=ssh-guide" rel="noopener noreferrer"&gt;helicon.sh&lt;/a&gt;. SSH projects were contributed by &lt;a href="https://github.com/dinhvh" rel="noopener noreferrer"&gt;Hoà Dinh&lt;/a&gt; in &lt;a href="https://github.com/HarjjotSinghh/helicon/pull/56" rel="noopener noreferrer"&gt;#56&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Helicon is an unofficial community project, not affiliated with Meta.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>ssh</category>
      <category>devtools</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I built a desktop app for Meta's Muse Code CLI, and the billing is the whole story</title>
      <dc:creator>Harjot Rana</dc:creator>
      <pubDate>Tue, 15 Sep 2026 23:25:08 +0000</pubDate>
      <link>https://dev.to/harjjotsinghh/i-built-a-desktop-app-for-metas-muse-code-cli-and-the-billing-is-the-whole-story-ane</link>
      <guid>https://dev.to/harjjotsinghh/i-built-a-desktop-app-for-metas-muse-code-cli-and-the-billing-is-the-whole-story-ane</guid>
      <description>&lt;p&gt;Meta shipped Muse Code in August: a coding agent that lives in your terminal, with a subscription starting at five dollars a month. The model is good. The harness runs subagents in parallel and keeps an append-only event log you can replay.&lt;/p&gt;

&lt;p&gt;Two things about it bothered me enough to spend a fortnight on them.&lt;/p&gt;

&lt;p&gt;It is terminal-only. And on Windows there is no native build at all, so several comparison guides now tell readers that if they are on Windows or want a desktop app, they should pick a competitor.&lt;/p&gt;

&lt;p&gt;So I built &lt;a href="https://github.com/HarjjotSinghh/helicon" rel="noopener noreferrer"&gt;Helicon&lt;/a&gt;: an open-source desktop and web client for the actual Muse CLI. This post is about the one design decision that shaped everything else, because it is the part I got wrong first and the part most people get wrong when they wrap a CLI.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mistake I made first: scraping the terminal
&lt;/h2&gt;

&lt;p&gt;The obvious way to put a UI on a CLI is to spawn it in a pty, parse what comes out, and render that.&lt;/p&gt;

&lt;p&gt;I did this. It took an afternoon and it worked, in the way that a demo works.&lt;/p&gt;

&lt;p&gt;Then I tried to implement approvals. The agent wants to run &lt;code&gt;npm test&lt;/code&gt;, and the user needs to allow or reject it. In a scraped terminal, "the agent is asking for approval" is a string you pattern-match out of ANSI escape codes, and your answer is keystrokes you write back into a pty. You are inferring a security decision from formatted text. Every output tweak upstream becomes a correctness bug in your client, and some of those bugs run commands nobody approved.&lt;/p&gt;

&lt;p&gt;I deleted it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What replaced it
&lt;/h2&gt;

&lt;p&gt;Meta publishes the &lt;a href="https://github.com/meta-models/muse-code-sdk" rel="noopener noreferrer"&gt;Muse Code SDK&lt;/a&gt;, MIT licensed, for programmatically driving Muse over the Muse Session Protocol. The CLI itself exposes &lt;code&gt;muse serve&lt;/code&gt;, which speaks that protocol over stdio.&lt;/p&gt;

&lt;p&gt;So Helicon's daemon spawns one &lt;code&gt;muse serve&lt;/code&gt; host per workspace and talks MSP to it through the official SDK. Muse is the agent. Helicon is a client.&lt;/p&gt;

&lt;p&gt;That means approvals arrive as protocol events, not as parsed text. Diffs arrive as structured data. Sessions that were started in the terminal show up in the sidebar with their history, because the CLI already wrote them to disk in a format the protocol can read back. Nothing is inferred from what the screen happened to look like.&lt;/p&gt;

&lt;p&gt;The rule I would give anyone wrapping a CLI: if the vendor documents a protocol, the protocol is the only surface that stays where you left it. Output formatting is not an API, no matter how stable it looks today.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that decides your bill
&lt;/h2&gt;

&lt;p&gt;Here is the thing I did not expect to be the most important feature.&lt;/p&gt;

&lt;p&gt;Muse Code subscriptions only bill through Muse's own harness. If you point a generic OpenAI-compatible client at the API, you are on pay-as-you-go rates instead. Same model, different bill.&lt;/p&gt;

&lt;p&gt;I have watched people work this out the hard way in public and reach the wrong conclusion. One thread has a user stating, twice, that a Muse Code subscription "can not be used in any kind of GUI harness" and that you have to go pay-as-you-go to use it anywhere but the terminal. Another person in the same thread confirms the symptom: using the vendor's harness deducts from the plan, and every other option lands on the API.&lt;/p&gt;

&lt;p&gt;The observation is correct. The conclusion is not.&lt;/p&gt;

&lt;p&gt;The subscription is not terminal-locked. It is harness-locked. Once you stop trying to replace the harness and start driving it instead, a GUI costs nothing extra, because the work is still going through Muse with your own &lt;code&gt;muse login&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That is the entire argument for this architecture, and it is worth more than any UI feature I could build. If your wrapper reimplements the agent loop against the raw model API, you have quietly moved your users onto a second bill for a model they already pay for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Windows, honestly
&lt;/h2&gt;

&lt;p&gt;Muse has no Windows binary. That is not something a client can fix.&lt;/p&gt;

&lt;p&gt;What Helicon does is more boring than "Muse Code on Windows" sounds. The daemon runs natively on Windows, and when it needs the agent it runs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;wsl &lt;span class="nt"&gt;-d&lt;/span&gt; Ubuntu &lt;span class="nt"&gt;--&lt;/span&gt; muse serve
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then it translates paths in both directions, because the daemon thinks in &lt;code&gt;C:\Users\you\project&lt;/code&gt; and the agent thinks in &lt;code&gt;/mnt/c/Users/you/project&lt;/code&gt;. Get that wrong and every file operation lands somewhere plausible and wrong.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;toWindowsPath&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;wslPath&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;match&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;wslPath&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;match&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/^&lt;/span&gt;&lt;span class="se"&gt;\/&lt;/span&gt;&lt;span class="sr"&gt;mnt&lt;/span&gt;&lt;span class="se"&gt;\/([&lt;/span&gt;&lt;span class="sr"&gt;a-z&lt;/span&gt;&lt;span class="se"&gt;])\/(&lt;/span&gt;&lt;span class="sr"&gt;.*&lt;/span&gt;&lt;span class="se"&gt;)&lt;/span&gt;&lt;span class="sr"&gt;$/&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;match&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`Cannot map to Windows: not a /mnt/&amp;lt;drive&amp;gt; path: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;wslPath&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;.`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;drive&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;match&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;toUpperCase&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;rest&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;match&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="se"&gt;\/&lt;/span&gt;&lt;span class="sr"&gt;/g&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;drive&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:&lt;/span&gt;&lt;span class="se"&gt;\\&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;rest&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note the throw. An unmappable path is a bug, not a thing to guess at.&lt;/p&gt;

&lt;p&gt;So WSL2 is still required, and I say that on the landing page, in the install steps and in the FAQ. What changes is not the requirement. It is that you stop living in an Ubuntu terminal on your own machine to use a tool you pay for. The installer is signed and auto-updates, which on Windows matters more than people from other platforms expect.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bundling the runtime
&lt;/h2&gt;

&lt;p&gt;The daemon is Node. Until last night, that meant a user had to install Node 22 or newer before Helicon would start, and on Windows they had to install it on the Windows side rather than inside WSL, which is exactly the kind of instruction that loses people.&lt;/p&gt;

&lt;p&gt;The fix is a Tauri sidecar. A build script downloads a pinned Node build, verifies its SHA256 against the published checksums, and drops it in &lt;code&gt;src-tauri/binaries/node-&amp;lt;target-triple&amp;gt;&lt;/code&gt;. For universal macOS builds, &lt;code&gt;lipo&lt;/code&gt; merges the two architectures. At boot the app prefers the binary next to its own executable and only falls back to a system Node if that is missing.&lt;/p&gt;

&lt;p&gt;It costs about 30MB on the Windows installer. It removes an entire step from the install instructions, and an entire category of "it doesn't launch" issue. Worth it.&lt;/p&gt;

&lt;p&gt;One detail worth copying: the bundled runtime's folder is deliberately left off the PATH handed to subprocesses. Users run shell commands through the app, and their &lt;code&gt;node&lt;/code&gt; should be their &lt;code&gt;node&lt;/code&gt;, not mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it looks like
&lt;/h2&gt;

&lt;p&gt;Projects grouped by folder, worktrees included. Threads with inline diffs where the edit happened. Approvals surfaced the moment they arrive, never batched, never bypassed. A cost view that shows what each thread would have cost at published API rates, so "is this plan worth it" becomes a number instead of a feeling. Command palette, slash commands, full keyboard operation.&lt;/p&gt;

&lt;p&gt;The same React UI ships twice: as a Tauri desktop app, and as a web app pointed at a remote daemon. One codebase, two shells.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest gaps
&lt;/h2&gt;

&lt;p&gt;The macOS builds are not Apple-notarized yet, so the first launch needs a right-click and Open. Linux has no packaged build, only a source path. The SDK is a developer preview, so the protocol can move under me. And Helicon is not the only GUI in this space — there are several, including ACP adapters for Zed and JetBrains, and an unofficial VS Code extension. If you want Muse inside your editor, use those. Helicon is for people who want a standalone workspace.&lt;/p&gt;

&lt;p&gt;It is free, MIT licensed, and unofficial: not made, sponsored or endorsed by Meta.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/HarjjotSinghh/helicon" rel="noopener noreferrer"&gt;github.com/HarjjotSinghh/helicon&lt;/a&gt; · &lt;a href="https://helicon.sh" rel="noopener noreferrer"&gt;helicon.sh&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you use Muse Code on Windows, I would like to know what breaks. That is the platform I have tested most and can verify least.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>tauri</category>
      <category>ai</category>
      <category>windows</category>
    </item>
    <item>
      <title>I mined 1,398 corrections out of my coding agent logs and turned them into a rules file</title>
      <dc:creator>Harjot Rana</dc:creator>
      <pubDate>Thu, 10 Sep 2026 13:08:17 +0000</pubDate>
      <link>https://dev.to/harjjotsinghh/i-mined-1398-corrections-out-of-my-coding-agent-logs-and-turned-them-into-a-rules-file-db9</link>
      <guid>https://dev.to/harjjotsinghh/i-mined-1398-corrections-out-of-my-coding-agent-logs-and-turned-them-into-a-rules-file-db9</guid>
      <description>&lt;p&gt;My coding agent asks me the same things constantly. New file or add to the existing one. Do you want a test for this. Should I clean up the function next to the one I am touching.&lt;/p&gt;

&lt;p&gt;I have answered all of these. Dozens of times. The answers were sitting in my chat history and nobody, me included, was reading them back.&lt;/p&gt;

&lt;p&gt;So I wrote a tool that does.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it actually is
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;jot&lt;/code&gt; is an extractor plus a knowledge base. The extractor finds the session stores your coding agents leave on disk, reads them, and pulls out the moments where you stopped the agent and said no, do it this way instead. Those get clustered into rules, each with a reason and a stated exception, and written to markdown.&lt;/p&gt;

&lt;p&gt;You load that markdown as an agent skill. Your agent reads it before it starts working.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add HarjjotSinghh/jot &lt;span class="nt"&gt;-g&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then in any agent that supports Agent Skills:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/jot should this state be global or local?
/jot review this PR
/jot which model should these subagents run on?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That command installs mine, which is really a worked example. The part worth your time is pointing the extractor at your own logs.&lt;/p&gt;

&lt;p&gt;No training. No GPU. No fine-tune. Text files.&lt;/p&gt;

&lt;h2&gt;
  
  
  The version I sat on for a year
&lt;/h2&gt;

&lt;p&gt;For about a year I wanted to fine-tune a small model, 8B to 20B, on every file I own. Text, audio, PDFs, spreadsheets, decks. A model of me, trained on me.&lt;/p&gt;

&lt;p&gt;I never wrote a line of it. Not once did I open a terminal.&lt;/p&gt;

&lt;p&gt;Then I read Kun Chen's post about distilling himself into a skill, and the thing that clicked was that the hard part of "clone yourself" was never the weights. It was the preferences. And preferences are small enough to just write down.&lt;/p&gt;

&lt;p&gt;By preferences I mean the boring calls with no correct answer, only my answer. Refactor now or ship and come back. Test first or backfill before merge. Ask before touching anything with real users on it, or just go. A generic model has a sensible default for every one of these. Several of mine are not the default, and that gap is the entire reason the agent keeps interrupting me.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do not interview yourself
&lt;/h2&gt;

&lt;p&gt;The obvious way to build this is to sit with a model and answer questions about how you work. It produces beautiful, useless output, because everyone is aspirational about themselves.&lt;/p&gt;

&lt;p&gt;Ask me if I write tests first and I will say yes with a straight face. Look at what I actually shipped and you find spikes with no tests and coverage backfilled right before merge. The second version is the true one.&lt;/p&gt;

&lt;p&gt;So I skipped the interview and went to the logs.&lt;/p&gt;

&lt;p&gt;The highest-value thing on my disk is not my code. It is every time I typed something at an agent that was one keystroke away from doing the wrong thing. Every agent saves those. Nobody reads them back.&lt;/p&gt;

&lt;h2&gt;
  
  
  The extraction
&lt;/h2&gt;

&lt;p&gt;Six adapters: Claude Code JSONL, Codex rollouts, Grok history, Gemini CLI logs, OpenCode, and Cursor, which stores every message as a row in a 1.5 GB SQLite file.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python scripts/extract_corpus.py
python scripts/extract_corpus.py &lt;span class="nt"&gt;--agents&lt;/span&gt; claude codex cursor &lt;span class="nt"&gt;--min-score&lt;/span&gt; 5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The pipeline:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Normalize everything to turns. Six agents, six schemas, one shape.&lt;/li&gt;
&lt;li&gt;Pair each thing I said with the agent turn it was reacting to. "No, not like that" is meaningless on its own. It only carries information next to the thing it rejected.&lt;/li&gt;
&lt;li&gt;Score against a lexicon weighted toward corrections, rejections, and stated preferences.&lt;/li&gt;
&lt;li&gt;Dedupe hard, because I complain about the same three things constantly and 40 instances of one gripe is one rule, not 40.
First pass: 1,398 distinct judgment events across 1,290 conversations. 505 corrections. 380 stated preferences.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Every rule in the knowledge files carries an evidence tag pointing back at the source:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;[observed Nx]&lt;/code&gt; extracted from N distinct real sessions, dated&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;[stated]&lt;/code&gt; I wrote it down as a rule&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;[inferred]&lt;/code&gt; derived from adjacent behaviour, lower confidence, flagged at use time
A rule with no evidence is a vibe. If it cannot be traced, it does not go in.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Layout
&lt;/h2&gt;

&lt;p&gt;A thin skill stub, and a knowledge base that lives in the repo.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;skills/jot/SKILL.md   loader; local checkout first, raw repo URLs second
ENTRY.md              routing table plus the decision procedure
PRINCIPLES.md         durable rules that survive a framework change
ENGINEERING.md        backend, debugging, git, review
FRONTEND.md           visual taste, design parity, slop detectors
AGENTIC.md            model routing, autonomy grants, parallelism
WORKFLOWS.md          named sequences: feature, bug, refactor, review, EOD
TOOLS.md              what I reach for, and the friction I hit
BOUNDARIES.md         hard stops, read before anything outward-facing
VOICE.md              how I write; loaded only for text going out under my name
OPERATING.md          attention, cadence, escalation
CONTEXT.md            who I am, what is active
state/                mined evidence, changelog, pending promotions
bench/                the benchmark that says whether any of this works
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The stub never changes. All churn happens in the knowledge files, so every agent pointed at the repo stays in sync without a separate update step.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;private/&lt;/code&gt; and &lt;code&gt;state/evidence/&lt;/code&gt; are gitignored and never fetched over the network. The public layer holds preferences. The private layer holds the specifics that make them actionable: named colleagues, clients, money, positioning. The skill only loads those from a local checkout.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does it work
&lt;/h2&gt;

&lt;p&gt;A rules file that sounds like you but decides differently is worse than no file at all. It makes confident wrong calls in your name and you are not there to catch them. So &lt;code&gt;bench/&lt;/code&gt; answers the question properly.&lt;/p&gt;

&lt;p&gt;Three arms. A control model with no skill. The same model with &lt;code&gt;/jot&lt;/code&gt; loaded. And me, answering blind.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python scripts/bench_run.py &lt;span class="nt"&gt;--set&lt;/span&gt; v1
python scripts/bench_score.py &lt;span class="nt"&gt;--run&lt;/span&gt; &amp;lt;timestamp&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ten questions, each one a decision where a competent generic model has a plausible default that is not what I actually do. Both model arms run in clean-room sessions that never saw the conversation where the rules were written, because a session that helped write the answers already knows the answers. Their responses go into sealed files I do not open until I have committed to mine.&lt;/p&gt;

&lt;p&gt;Scored 0 to 2 against my answer. Minus one for fabricating a preference I do not hold.&lt;/p&gt;

&lt;p&gt;First run: skill 13/20, control 7/20.&lt;/p&gt;

&lt;p&gt;The number that matters is the delta between the two model arms. A skill that scores well only because any competent model would have scored well is recording things that did not need recording.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then I sorted the results by where each rule came from
&lt;/h2&gt;

&lt;p&gt;I had written two kinds of rules without really noticing. Ones mined out of real transcripts, and ones I had simply believed about myself and typed in.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;mined from transcripts:  2, 2, 2, 2
believed about myself:   0, 1, 0, 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No exceptions in either direction. The mined rules carried the entire gain. Everything I had assumed about myself was dead weight or actively wrong, and it was wrong in exactly the same confident tone as the stuff that was right.&lt;/p&gt;

&lt;p&gt;The worst one: my file said that once I give a blanket approval, the agent should stop asking. I had generalized that from exactly one line I typed mid-task, once. The benchmark asked what to do at a sub-decision that approval never covered. The skill told the agent not to ask. My real answer was "I would ask first."&lt;/p&gt;

&lt;p&gt;One occurrence is an anecdote. My own precedence rules say three occurrences make a rule. I broke my own precedence rules inside the file that contains them.&lt;/p&gt;

&lt;p&gt;That is how this goes wrong. Not by being wrong about you. By being right about you in one situation and then applying it everywhere.&lt;/p&gt;

&lt;h2&gt;
  
  
  Patched four rules, re-ran, 18/20
&lt;/h2&gt;

&lt;p&gt;And one question that had scored 2 dropped to 0.&lt;/p&gt;

&lt;p&gt;The rule I added to fix "just take the two-hour refactor now" got applied to a two-minute fix on a live customer list. Which is the opposite of what I do, because that is not my code and those are somebody's real customers.&lt;/p&gt;

&lt;p&gt;A rule that fixes one case and breaks another is written too broadly. Misses now become golden cases in &lt;code&gt;bench/golden/&lt;/code&gt;, and each one turns into a rule with an explicit &lt;code&gt;Doesn't apply when&lt;/code&gt; line. Boundaries outrank every principle in the repo instead of politely tie-breaking with them.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it is like to use
&lt;/h2&gt;

&lt;p&gt;Uncanny. Not "wow, impressive model" uncanny. Specifically: you ask it something you have never thought about before, it answers, and you get a small jolt of yes, that is exactly what I would have done, and that is exactly why.&lt;/p&gt;

&lt;p&gt;It is talking to yourself, except this version of you has read every correction you have ever issued and never gets tired.&lt;/p&gt;

&lt;p&gt;It disagrees sometimes, and the disagreements turned out to be more useful than the agreements. It was never being stupid. It was faithfully applying a rule I had written badly. The gap between the two model arms measures my preferences. The gap between the skill and me measures my writing. Two different problems, fixable separately.&lt;/p&gt;

&lt;p&gt;Worth doing even if nobody else installs it. Writing a rule with a reason, a stated exception and a piece of evidence forces you to find out whether you actually hold it. A few of mine did not survive the format.&lt;/p&gt;

&lt;p&gt;Not bad for an idea I sat on for a year because I thought it needed a GPU.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx skills add HarjjotSinghh/jot &lt;span class="nt"&gt;-g&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Repo: &lt;a href="https://github.com/HarjjotSinghh/jot" rel="noopener noreferrer"&gt;https://github.com/HarjjotSinghh/jot&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Prior art: &lt;a href="https://github.com/kunchenguid/kun" rel="noopener noreferrer"&gt;https://github.com/kunchenguid/kun&lt;/a&gt;, which made the case that the repeatable part of your judgment is worth externalising, and that it was never the part that was your moat. Read Kun's first.&lt;/p&gt;

&lt;p&gt;If you run the extractor on your own logs, I want to know two things. How many events it found, and how many of the rules you wrote by hand survived contact with a benchmark. My guess is fewer than you expect.&lt;/p&gt;

</description>
      <category>showdev</category>
      <category>ai</category>
      <category>productivity</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How we built a 14-agent pipeline that ships a deployed app + launch assets in ~7 minutes</title>
      <dc:creator>Harjot Rana</dc:creator>
      <pubDate>Mon, 01 Jun 2026 01:54:28 +0000</pubDate>
      <link>https://dev.to/harjjotsinghh/how-we-built-a-14-agent-pipeline-that-ships-a-deployed-app-launch-assets-in-7-minutes-24ko</link>
      <guid>https://dev.to/harjjotsinghh/how-we-built-a-14-agent-pipeline-that-ships-a-deployed-app-launch-assets-in-7-minutes-24ko</guid>
      <description>&lt;p&gt;Most AI app builders stop at "deployed." You prompt, you get a repo, maybe a preview URL, and then the actual work starts: wiring a domain, writing the landing copy, cutting screenshots, drafting the launch thread. We wanted the pipeline to stop at "launched" instead, so we built one. This is how it works under the hood, including the parts that broke.&lt;/p&gt;

&lt;p&gt;The product is &lt;a href="https://moonshift.io" rel="noopener noreferrer"&gt;Moonshift&lt;/a&gt;. One prompt triggers 14 specialized agents across 10 phases. Average run is ~7 minutes and ~$3 in API spend, with a hard $5 ceiling that aborts the run. Everything ships to your Vercel, your GitHub, your database. This post is the engineering, not the pitch.&lt;/p&gt;

&lt;h2&gt;
  
  
  The core problem: parallel agents drift
&lt;/h2&gt;

&lt;p&gt;The naive version of "many agents build an app" falls apart fast. If a backend agent and a frontend agent both work from a vague English spec, they invent incompatible contracts. The backend returns &lt;code&gt;{ user_id }&lt;/code&gt;, the frontend reads &lt;code&gt;userId&lt;/code&gt;, and you find out at runtime in production.&lt;/p&gt;

&lt;p&gt;Our fix is a &lt;strong&gt;planner that emits a JSON contract first&lt;/strong&gt;. Before any code is written, one agent produces a typed contract: routes, request/response shapes, table schemas, env vars, page list. That contract is the single source of truth. Backend, frontend, database, and test agents all build against it in parallel instead of against prose.&lt;/p&gt;

&lt;p&gt;Then a &lt;strong&gt;contract-validator agent&lt;/strong&gt; runs after the parallel build and diffs the actual code against the contract. When the frontend's fetch shape doesn't match the backend's handler, the validator doesn't just flag it. It patches the mismatch. This one agent removed the largest single class of "looks done, 500s on click" failures we had.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 10 phases
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Plan&lt;/strong&gt; - generate the JSON contract.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scaffold&lt;/strong&gt; - lay down the framework skeleton (Next.js, config, deny-globs that protect files agents shouldn't touch).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backend&lt;/strong&gt; - API routes and server logic against the contract.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Frontend&lt;/strong&gt; - pages and components against the same contract.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Database&lt;/strong&gt; - schema + migrations (Drizzle + Turso).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validate&lt;/strong&gt; - contract-validator reconciles 3-5 in parallel, auto-fixes drift.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test + fix&lt;/strong&gt; - generated tests run; a fixer loop addresses failures.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deploy&lt;/strong&gt; - ships to &lt;em&gt;your&lt;/em&gt; Vercel via your token.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit&lt;/strong&gt; - security and a11y passes on the live deployment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Market + publish&lt;/strong&gt; - a marketer agent drafts X and LinkedIn launch posts in your voice, image-gen produces hero images, and a publisher gates everything behind your one-tap approval.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The interesting phases are 6, 7, and 10. Everyone has a code-gen step. Almost nobody has a reconcile step, a real fixer loop, or a phase whose only job is launch assets.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reliability: not every failure is equal
&lt;/h2&gt;

&lt;p&gt;Long multi-agent runs fail in boring ways: a rate limit, a flaky deploy, a model that returns prose where you asked for JSON. If you retry all of them the same way, you either give up too early on transient errors or burn money death-looping on deterministic ones.&lt;/p&gt;

&lt;p&gt;We run a &lt;strong&gt;failure classifier&lt;/strong&gt; that buckets every failure into &lt;code&gt;transient&lt;/code&gt;, &lt;code&gt;deterministic&lt;/code&gt;, or &lt;code&gt;permanent&lt;/code&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;transient&lt;/strong&gt; (429s, network blips, stream idle) - retry with backoff.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;deterministic&lt;/strong&gt; (a test that fails the same way every time) - hand to a fixer agent, don't blindly retry the same call.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;permanent&lt;/strong&gt; (bad auth, missing token) - stop and surface it. No point spending more.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Retries are capped on three axes at once: per-phase, per-agent, and a global per-run ceiling. The global cap is what keeps a single bad run from quietly turning into a $40 bill. Combined with the hard $5 abort, the worst case is bounded and visible instead of a surprise invoice.&lt;/p&gt;

&lt;p&gt;A subtle one we hit: a long LLM stream can go &lt;em&gt;idle&lt;/em&gt; without erroring (the upstream connection gets severed but the socket never closes). A naive loop waits forever. We added an idle watchdog in the agent loop so a silent stall is treated as a transient failure and retried, instead of hanging the whole run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Design constraints that shaped everything
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your infra, zero lock-in.&lt;/strong&gt; Code lands in your GitHub, the app deploys to your Vercel, the database is yours. Cancel the subscription and you keep a working product. This forced the deployer to operate purely through user-supplied tokens, which is more work than deploying to our own infra but is the entire point.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The publisher physically cannot post without you.&lt;/strong&gt; Social publishing is gated per post, per platform, behind an explicit human tap. Autonomy ends at the point where it would speak as you in public. Generation is automatic; publishing is not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hard cost ceiling.&lt;/strong&gt; $5/run, enforced mid-run, not reconciled after. Agents check remaining budget before expensive calls.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Stack
&lt;/h2&gt;

&lt;p&gt;Next.js for web, Drizzle + Turso (libSQL) for data, Playwright for browser automation in the marketing/audit phases. The orchestrator is a separate runtime from the web app, spawned per run from source so a fix ships without a full rebuild.&lt;/p&gt;

&lt;h2&gt;
  
  
  Honest lessons
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A typed contract beats a smarter prompt.&lt;/strong&gt; We spent weeks trying to make agents "just agree." Making them agree on a machine-checkable artifact was the actual fix.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A reconcile phase is worth more than a better code-gen model.&lt;/strong&gt; Catching drift after the fact, cheaply, beat every attempt to prevent it perfectly up front.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Classify failures before you retry them.&lt;/strong&gt; Uniform retry is how multi-agent systems burn money and still fail.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bound the blast radius in money, not just time.&lt;/strong&gt; A per-run dollar cap is the single most important guardrail in an autonomous pipeline that calls paid APIs in a loop.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you want to see the output, the first run is free with your own API key at &lt;a href="https://moonshift.io" rel="noopener noreferrer"&gt;moonshift.io&lt;/a&gt;. Happy to answer architecture questions in the comments.&lt;/p&gt;

</description>
      <category>devtools</category>
    </item>
  </channel>
</rss>
