Someone in a founder group I am part of asked a good question: how do you let several agents work on one product in parallel without them stepping on each other, and without the token bill getting out of hand?
I have been doing exactly that for a while. I build DuctTape.io, a diagram editor for AI architectures, as a single person. The rest of the team are Claude Code agents. This post is the honest version of how that works: the setup, the rules, what broke, and what I would tell anyone starting today.
The team
Seven Claude Code sessions run in parallel on one server. Each is a long-running session with a name, a role and a clear area it owns. They do not start fresh for every task; they stay open for days and keep their context.
- Product owner. Decides what gets built and in which order, and answers product questions from the others.
- Requirements. Writes every decision into the spec, a living document that is the single source of truth for scope.
- Developer. The only session that changes the web app code. Writes the code and the tests, runs the test suite and ships.
- Desktop developer. Builds a desktop app for Windows, macOS and Linux in its own repository. The newest role: it came in when a clearly separate second job appeared.
- Marketing. Website copy, e-mails, posts, articles, videos. Prepares everything, publishes nothing on its own.
- UX review. Clicks through the real build, measures, and files findings with a number.
- Customer success. Reads user feedback, reports and requests. Read-only: it can look, it cannot change anything.
And then there is me. I decide.
The direction of work is simple. I talk to the product owner. The product owner sets priorities and hands work to the others. Requirements writes it down, the developers build it, UX review measures it, marketing tells the world. Customer success brings what users say back to the product owner.
How they work together
Three things hold it together.
Messages between sessions. The sessions can send each other short messages. "This is live." "Please write the texts for the contact form, German and English." "The test server is busy for the next 40 minutes, please do not record." Most coordination happens this way, without me.
Documents as the source of truth. The spec, a UX review document with every finding and its status, and a customer success document. Agents forget; documents do not. When a session's context runs out, it reads the documents again and continues.
Project rules and memory. A CLAUDE.md in the repository holds the rules every session follows: wording ("component", never "block"), how to commit, which folders belong to whom, which conventions differ from what the model learned in training. Each project also has a small memory of decisions and preferences that survive across sessions.
Here is how one piece of work typically travels, from my idea to a live feature:
A recent example: I wanted a third kind of animation on connections in diagrams, for files and uploads. I told the product owner. The product owner asked marketing for the name and the one-line explanation, and requirements for the spec entry. Marketing proposed "Document" and which example architectures should use it; the product owner approved; requirements wrote it down. UX review raised three objections, the product owner settled them, and it is now in the developer's queue. I saw the decisions, not the twenty messages in between.
Where a human has to say yes
The agents can do almost everything. There are a few things they are not allowed to do on their own, and these rules matter more than any prompt:
- Anything public. Posts on LinkedIn, X or DEV, a published article, a video upload. Agents write the draft; I press the button.
- Releases. Agents may prepare a release draft with the notes. Publishing it is mine.
- Production data. Deleting or changing production data needs my explicit yes.
- Money. Nothing that costs money happens without me.
- Secrets. Keys and tokens are mine. An agent that needs one asks; it does not go looking.
One rule turned out to be important: a "yes" from one agent to another only counts when it is a real, marked approval. The product owner may release work inside the team, and it writes "approval" and what exactly it approves. Anything public still comes to me. An agent cannot know what I actually said, so a casual "the founder agreed" in a message is not enough.
What went wrong
Flaky end-to-end tests. For days the test suite failed at random. Clicks that missed, connections that reset, a dev server that died mid-run. The agents tried retries, longer timeouts, smaller test batches. The real cause was boring: the server was too small. Two cores and 8 GB of memory, shared by six sessions, a production build and a headless browser. The kernel killed processes when memory ran out. Full runs took 30 to 40 minutes when they finished at all; the worst one hung for more than 70 minutes. After upgrading to four cores and 16 GB, the chains of timeouts were gone. A full run takes about as long as before, but every run finishes.
What remained were three real test bugs. We tried two parallel workers on the bigger machine, and the run dropped to about 15 minutes, but two tests turned out to share one account. On the small machine they had never run at the same time. A faster machine does not only make tests faster, it also shows you which ones depended on being slow. We went back to one worker until those tests are fixed.
The lesson: when agents keep fixing symptoms, check the machine they run on.
Messages that did not arrive. Some sends between sessions failed with an error, one message was dropped by the receiver as a duplicate, and one session was not reachable at all for a while. Work then waited for an answer nobody had seen. Now every important handover gets a short confirmation back, and the shared documents record decisions, so nothing depends on a single message.
"Done" that was not done. An agent will happily report a task as finished before checking the result. The marketing agent once reported a note in the plan as written before it had written it, and the product owner once told another agent a message had been delivered before checking; the send had failed. My rule since then: show me proof. A screenshot, a test run, a link that works, a measured value. The UX review session exists partly for this reason: it measures the real build, not the developer's description of it.
Too much in one turn. Long sessions accumulate context. When it runs out, the session is summarised and continues, and small details can get lost in the summary. Writing the state down in the documents, not only in the conversation, solved most of this.
What about the cost?
This was the second half of the question, so here is what I actually do.
A subscription, not pay-per-token. All sessions run on a Claude subscription with a fixed monthly price. That makes the cost predictable: I pay the same whether the agents have a quiet day or a busy one. The limit is usage within the plan, not a growing invoice.
Different models for different roles. Not every role needs the strongest model. Each session runs on the model that fits its work: the most capable where the work is hardest, a lighter one where it is routine.
Seven sessions open, two or three working. All seven sessions stay open, but most of the time only two or three are actively working. The others are waiting: for a message, for a build to finish, for my approval. An idle session costs almost nothing. The approval points act as a natural brake, because nothing public moves until I have looked at it.
Documents instead of long context. Agents that re-read a focused document need less context than agents that drag a whole history around. Clear ownership also prevents duplicate work: two sessions never build the same thing.
What I would tell someone starting
- Give every agent one area it owns, and write it down. Most coordination problems are ownership problems.
- Keep the source of truth outside the chat. A spec and a few living documents beat any amount of context.
- Define the approval points before the first agent writes anything. Public, money, production data, secrets.
- Ask for proof, not for status. "Done" means a screenshot, a test run or a link.
- Size the machine for the agents plus everything they run. Builds, browsers and tests need memory too.
- Start with fewer agents than you think. Add a role when one session clearly has two jobs.
It is not magic, and it is not autonomous. It is a small team with very clear rules, where one person decides and the agents do most of the work.
If you want to see what this team builds: DuctTape.io is a diagram editor for AI architectures, and your own agents can draw in it over MCP.




Top comments (0)