The interesting decision in munder-difflin is what it keeps. The coding CLIs you already use stay at the center. It launches the coding CLIs you already have installed, claude, codex, agy, grok, kimi, qwen, opencode, crush, pi, copilot or a custom command, as real processes, and then builds an office around them. The README puts the economics plainly: it "works with the subscriptions you already pay for, on their hourly limits." Every agent in the office is one of those terminals, running with your existing subscription and its hourly limits. The framing is working inside those limits, and the README does not pitch it as a way past them.
The architecture described in the README starts with that terminal model, so it is worth following through.
A terminal is the unit of work
Each agent is a normal terminal process spawned in a pseudo-terminal through node-pty and rendered with xterm.js. The README calls the result "byte-for-byte authentic," meaning the harness shows you the same session the CLI would show you on its own. Every agent gets its own working directory, identity and provider-specific lifecycle.
The README gives provider support a simple rule: if it runs in a terminal, it can run here. The supported list covers Claude Code, Codex, Grok, Kimi Code, Gemini CLI, Antigravity, Qwen, OpenCode, Crush, Pi, GitHub Copilot and Cursor. It also supports your own API keys and local models through Ollama, LM Studio or vLLM.
The practical upside is that you keep the tool you already know. Click a desk on the floor and you read that terminal live, and you can type straight back into it.
The hive is a git repo of plain files
Once you have a dozen independent processes, something has to let them talk. Here that something is deliberately boring: a local git repo of plain files. Each agent writes to its own outbox/, and the harness's router delivers messages into the recipients' inbox/ folders. Agents read their memory and drain their mailbox.
The README also notes that no agent touches git. The harness is the single committer, a design it describes as avoiding index.lock corruption.
Memory follows the same file-first pattern. Each agent keeps markdown memory, which gets mined into a shared, searchable store the README calls a palace, backed by a semantic recall index. The stated goal is that you can close the app, come back the next day, and the agents still know what they learned.
One agent you brief
Running many subscriptions in parallel creates a new problem, which is that you now have many conversations to manage. The answer is an orchestrator the README calls your clone, named Michael. You brief Michael, and the GOD agent layer handles the roster, routing, a blackboard and a task ledger.
The supervisor resolves routine requests itself and escalates only critical items into an approvals queue you act on. The README lists those as spend, destructive operations and scope changes. Per-agent autonomy settings control how far each agent may go alone, and a circuit breaker steers, constrains and then stops anything that loops or runs away. For an office of processes spending your hourly allowance, a stop condition for loops is the feature I would look at first.
The office floor and the setup path
The visual layer is a Pixi.js floor where each session appears as an avatar. Agents walk to stations as they work, and envelopes fly between desks when they message each other. The moving avatars and envelopes give a visual read on which agents are working and which are messaging each other.
Adding an agent means picking the CLI, the model and the autonomy level, then assigning a desk. There is an Agent Gallery of ready-made roles to import. A first-run wizard checks what you already have installed and offers to install what is missing.
The app is built on Electron, React and TypeScript, runs on macOS, Windows and Linux, and is MIT licensed. The README marks it as pre-release at version 0.4.6. macOS builds are signed and notarized, so you do not need to build from source to try it.
If your team already pays for one or two of these CLIs and keeps wishing it could run several sessions on separate tasks without babysitting each one, this is a concrete design to evaluate. Start small, set autonomy conservatively, and observe how the agents coordinate within the hourly limits you already have.
GitHub: https://github.com/chaitanyagiri/munder-difflin
Curated by Agent Palisade — practical AI for small and mid-sized businesses.
Top comments (0)