DEV Community

Aviv Laufer
Aviv Laufer

Posted on

I talk to one agent. It runs the others.

The map of my setup: OMP in the middle, the tools it dispatches, and the 13 parts coming next.

Most of my coding day happens in one terminal pane. I type into OMP, and OMP does the rest.

It picks a model for each job. Planning goes to Opus. Normal turns go to Sonnet. Reviews go to a GPT model, so the reviewer is never the same model that wrote the code. When the work is big, OMP hands it to a fleet: Firstmate spawns crew agents, each one in its own worktree, and most of the time the crew is working and I'm not.

I use several coding agents, but I don't configure them separately. One repo decides which rules, skills, identities and models every agent gets, on every machine.

That repo is my dotfiles. It started in December 2016 with a tmux and vim config. In the last three months, since agents took over my workflow, it got almost as many commits as in the nine years before.

This post is the map: every tool, how they connect, what broke when I first wired them together, and the fix.

The map

                         me
                          |
                          v
   +---------------------------------------------+
   |  OMP  (my main tool, one model per role)    |
   +---------------------------------------------+
            |                        |
            | talks to models        | fm <project>
            | directly               v
            |              +----------------------+
            |              |  Firstmate           |
            |              |  (first mate on OMP) |
            |              +----------------------+
            |                 |   dispatch rules
            v                 v
   Anthropic, Cursor,     crew: OMP, Claude Code,
   OpenAI, local model    Codex, Grok
                              |
                              v
                     treehouse worktrees
                     (one per task, pooled)

   herdr on the laptop, tmux on the VM: every agent gets a pane
   dotfiles: renders the config for all of the above
Enter fullscreen mode Exit fullscreen mode

It runs on two laptops and a cloud VM. My work laptop gets the company rules: our skills, our ticket conventions, our compliance rules. My personal laptop should not see any of that, and the two use different model accounts.

I'll go through each box below. First, the part that went wrong.

What broke

OMP is where I type, but it is not the only agent that runs. Firstmate starts Claude Code, Codex or Grok as crew, and Cursor agent is there too. Each of these tools reads its instructions from a different place:

~/.claude/CLAUDE.md
~/.codex/AGENTS.md
~/.pi/agent/AGENTS.md
~/.agents/AGENTS.md
~/.gemini/GEMINI.md
~/.config/opencode/AGENTS.md
~/.copilot/copilot-instructions.md
Enter fullscreen mode Exit fullscreen mode

Seven files that should say the same thing. Skills have the same problem: each tool has its own skills folder.

Agents read their instructions in every session, so when those instructions drift, the agents drift with them. Nothing warns you.

Two things were wrong.

First, the files had drifted. My ~/.codex/AGENTS.md was a copy I'd made by hand in May, and it was out of date. Pi, another agent I had installed, saw only 3 of my 23 skills. Each tool was working with a different set of rules, and I had assumed they were all the same.

Second, the instructions were too long. The global file had grown to 465 lines, about 21 KB or 5,300 tokens. It loads on every turn, in Claude Code and each of its subagents, in OMP, Codex, Pi, Gemini, OpenCode and Copilot. So every one of those paid for all 465 lines, every time.

The fix

One source, compiled per machine

I split the instructions into small fragments, one topic per file:

agents/fragments/
  00-header.md
  05-no-emojis.md
  10-voice.md
  20-workflow.md
  30-git-conventions.md
  32-worktrees-stacking.md
  35-pull-requests.md
  45-core-principles.md
  ...
  50-work-skills.md     # work only
  55-compliance.md      # work only
Enter fullscreen mode Exit fullscreen mode

Each machine profile is a plain list of the fragments it gets. The personal profile leaves out the work-only ones:

# agents/machines/personal.txt
00-header.md
05-no-emojis.md
20-workflow.md
35-pull-requests.md
32-worktrees-stacking.md
45-core-principles.md
30-git-conventions.md
40-languages.md
10-voice.md
Enter fullscreen mode Exit fullscreen mode

make build puts them together into one flat AGENTS.md, with a marker before each fragment so I can see where a rule came from:

while IFS= read -r name; do
  case "$name" in ''|'#'*) continue ;; esac
  printf '<!-- Source: fragments/%s -->\n' "$name"
  cat "fragments/$name"
done < machines/$(PROFILE).txt
Enter fullscreen mode Exit fullscreen mode

Why flat and not includes? Codex, Pi and Cursor don't expand any include syntax. So the modular part lives in the source, and every tool gets one plain file.

make deploy then symlinks that one file into all seven paths. There are no copies left to go stale. When I edit a fragment and run make build, every tool picks it up in its next session.

A check that blocks drift

A pre-commit hook runs make check. It builds again into a temp folder and diffs the result against the built file. If I edit a fragment and forget to build, the commit fails:

check: build/personal/AGENTS.md is out of date. Run: make -C agents build
Enter fullscreen mode Exit fullscreen mode

One skills folder

Skills go in one place, ~/.agents/skills/. When I checked in July, Codex, Cursor, OpenCode, Gemini and Copilot all read that folder directly. Claude Code doesn't, so make skills-sync adds a symlink per skill into ~/.claude/skills/. It does the same for Pi's own folder, which is where the 3-of-23 problem was.

It never overwrites a real file. A skill installed some other way gets skipped with a warning, not deleted.

Then I could cut it in half

With one file, I could finally see everything in one place. So I cut it. 465 lines down to 215. About 5,300 tokens down to 2,500.

The biggest cut was the pull request procedure. I moved it, as is, into a skill. It loads only when an agent works on a PR. The fragment kept only the hard rules: draft only, no merge and no deploy without my approval. The rest of the fragments I trimmed down to the rules, without the explanations.

With seven copies, I never cut anything. I didn't know which copy was the real one, so I only added. With one source, cutting is safe.

Now I split instructions into two kinds. Rules every agent needs on every turn go in the global file, and I keep it short. Procedures for one kind of task go in a skill. An agent renaming a variable doesn't need my PR procedure, and it shouldn't pay for it in tokens.

The rest of the map

With that in place, here is what each box in the diagram does.

OMP

OMP (Oh My Pi) is a coding agent that talks to the model APIs directly. I log in to Anthropic, Cursor and OpenAI once, and OMP routes each kind of work to a model through roles:

modelRoles:
  default: cursor/claude-sonnet-5-5:medium
  plan:    cursor/claude-opus-5-5:medium
  task:    cursor/claude-sonnet-5-5:medium
  slow:    cursor/composer-2.5:high
  smol:    cursor/gpt-5.6-luna:medium
  commit:  cursor/gpt-5.6-luna:low
  tiny:    local/lfm2.5-230m
Enter fullscreen mode Exit fullscreen mode

The reviewer is set in a separate block of the same config file:

task:
  agentModelOverrides:
    reviewer: cursor/gpt-5.6-terra:high
Enter fullscreen mode Exit fullscreen mode

That's my personal laptop. It only has Claude and Cursor accounts, so almost everything goes through Cursor. The work laptop uses the same roles with different providers, and dotfiles picks the right ones per machine.

Model names change every few months, so here is the current table. I'll keep it up to date. The pattern is what matters.

Role What it does Current model
default normal turns Claude Sonnet 5.5
plan planning Claude Opus 5.5
task subagent work Claude Sonnet 5.5
slow hard fixes Composer 2.5
reviewer code review, different vendor than the author GPT-5.6 Terra
smol cheap fan-out GPT-5.6 Luna
commit commit messages GPT-5.6 Luna, low effort
tiny throwaway calls a 230M local model
judge "is this actually done?" a small model on Cloudflare Workers AI

Firstmate

Firstmate runs a fleet of agents. I'm the captain. The first mate is an OMP session that takes my ask, writes a brief and starts crew agents to do the work. fm <project> opens it for one project, and each project gets its own home, so backlogs and running tasks never mix.

The crew is OMP by default, but a few dispatch rules send work to other tools:

{
  "when": "The task is a trivial mechanical edit: rote rename, formatting sweep, targeted typo fix, dependency bump, or gathering files.",
  "use": { "harness": "claude", "model": "haiku", "effort": "low" }
},
{
  "when": "The task is a scout: investigate, audit, trace a bug, or produce a standalone report without changing code.",
  "use": { "harness": "omp", "effort": "high" }
}
Enter fullscreen mode Exit fullscreen mode

My preferences for the first mate live in a markdown file it reads at the start of every session. A few lines from it:

- Never use em dashes in PR titles, PR bodies, commit messages, or comments.
- Every PR any worker opens is created in draft mode, no exceptions.
- Escalate to me on any change touching auth, tenancy boundaries, data
  retention, or migration ordering, even when the tests are green.
- Scout tasks end in a report, never a code change.
Enter fullscreen mode Exit fullscreen mode

Worktrees: two tools on purpose

When I start an agent myself, I use worktrunk (wt). One alias creates a branch, a worktree and an agent session:

worktree-path = "~/worktrees/{{ repo }}/{{ branch | sanitize }}"

[aliases]
cc  = "wt switch --create {{ branch }} -x claude"
cur = "wt switch --create {{ branch }} -x wt-agent -- cursor-agent"
Enter fullscreen mode Exit fullscreen mode

When Firstmate starts a crew agent, it uses treehouse. Treehouse keeps a pool of worktrees per repo and leases one out per task (treehouse get --lease, then treehouse return). Two crew agents on the same repo get slot 1 and slot 2 and never collide. Most of the worktrees on my machine come from treehouse, because most of the time the crew is the one working.

herdr and tmux

herdr is the terminal multiplexer on my laptops. Unlike tmux, it knows which pane is an agent and whether that agent is working, blocked, done or idle:

[ui]
agent_panel_sort = "spaces"     # one space per repo/worktree
status_indicators = "symbols"   # blocked/working/done/idle
show_agent_labels_on_pane_borders = true
Enter fullscreen mode Exit fullscreen mode

With several crew agents running, ctrl+alt+up and ctrl+alt+down jump between them. On the cloud VM I use plain tmux, so agents keep running after I close the laptop.

Dotfiles

Everything above is rendered from the dotfiles repo, in layers:

base < os < organization < machine < local.toml
Enter fullscreen mode Exit fullscreen mode

Each layer only overrides the values it names. The work laptop declares the work organization and the personal one. The personal laptop declares only personal. A repo outside every declared root has no git identity at all, so a commit fails instead of going out under the wrong name.

Try it

The companion repo has a small working version of the fragments build: four fragments (one of them work-only), two profiles, a pre-commit hook, and the Makefile with build, check and deploy. By default it deploys into a fake home folder, so it never touches your real config. When you point it at your real home, it skips any real file in the way instead of overwriting it.

make build && make deploy
ls -la fake-home/.codex   # AGENTS.md -> .../build/personal/AGENTS.md
Enter fullscreen mode Exit fullscreen mode

golem-notes-examples/01-the-map

What's coming

One part every two weeks:

  1. The map (this post)
  2. Close the laptop, the agent keeps working: agents on a cloud VM with tmux
  3. My agents on two laptops, controlled from one phone app
  4. Every PR gets an agent review before a human sees it
  5. Captain, first mate, crew: orchestrating a fleet of agents
  6. Nine roles, not one model
  7. Right model for the task: dispatch rules and cheaper subagents
  8. When the credits run out: fallback chains
  9. A cheap judge for "is this actually done?"
  10. Teaching agents how I work: my skills library
  11. Dotfiles an agent can't break
  12. One laptop, many companies: identities and secrets
  13. Testing your dotfiles like production code
  14. What I'd do differently

If you run more than one agent tool, I'd like to know how you keep their rules in sync. Leave a comment.


This is part 1 of my series on running a real agentic dev setup. I publish each part first on Golem Notes, with the full write-up and what's coming next. Subscribe there to get the next one in your inbox: https://golemnotes.substack.com/p/i-talk-to-one-agent-it-runs-the-others?utm_source=devto&utm_medium=social&utm_campaign=agentic-setup-01

Top comments (0)