DEV Community

Cover image for The Day I Left the Team
Shinsuke KAGAWA
Shinsuke KAGAWA

Posted on Originally published at norsica.jp

The Day I Left the Team

The last thing I had open was the orchestrator's session. It was a Claude Code window I kept running so the team had someone to route work, and so I had someone to talk to. Today I closed it.

A few minutes later I asked, in Slack, how things were going. The reply came back in the thread about thirty seconds later: what had been agreed, what was being built, who was reviewing it, and what would happen next. It ended with one more line: nothing is needed from you right now.

That is a team I set up about three days ago. It is a small group of AI agents that decides what to make, makes it, reviews its own work, and reports to me. I am no longer on it. I am closer to a stakeholder: I say where I want the product to go, I answer when asked, and I can stop everything.

Who is on the team

There are six members, and none of them talk to each other directly. Everything goes through a board: a folder of markdown files, one request per file, each with a status that moves from open to doing to done, blocked, or failed.

A bash script called studio watches the board every thirty seconds and launches whoever an item is addressed to. It also launches the orchestrator whenever something needs routing, runs a patrol every thirty minutes while anything is being built, and carries messages between the board and my Slack channel. The orchestrator keeps the flow moving. Two directors, one on Claude and one on GPT, decide the direction. Three executors do the work: luna for regular implementation, sol for what luna cannot finish, and visual for images.

The plumbing is scripts on purpose. Waiting, launching, handing a discussion from one director to the other, marking a crashed run as failed: an LLM does not have to remember any of it. My first version had the orchestrator wait for the board itself, and it forgot to wait again after each pass.

How it decides

Direction is decided by the two directors reaching one conclusion together. Getting that to work took more rework than anything else in the system, because every version I tried failed in a way I only saw afterwards.

The first version had each director write a proposal without seeing the other's, then a cross-review: each read the other's proposal and chose again, at the same time, without seeing the other's new choice. A rule told them to yield when the difference was a matter of taste, so both yielded at once and simply traded places. Over the following rounds they kept switching to each other's side, four times in all. Nobody was arguing. Each round was two blind votes. I replaced the cross-review with a direct discussion, where each writes an opening and then they take turns in the same file until one proposes a conclusion and the other agrees.

That worked better, and it surfaced the next problem. Across six discussions, the director who took the first turn wrote the final conclusion in four, and three closed on the very first turn with the other side simply agreeing. First mover anchoring, in a team of two. Studio now picks the first turn from a hash of the item name instead of whoever finished their opening first, and a conclusion can only be proposed after both sides have written at least one turn.

The orchestrator was part of the problem too. It wrote the briefs for the directors, and in writing them it paraphrased me, framed the options, and occasionally suggested what to decide. It was acting like a project manager, and the directors took its brief as their frame. I rewrote its role as a scrum master: it keeps information moving and leaves what to do to the directors. A brief is now made of quotes only: the item that started the decision and my words, in full.

The next failure started with my own feedback. I told the directors that the product had drifted into looking like everything else of its kind, and that its original distinctiveness had become secondary. Both directors read it, agreed with it, and proposed a fix that kept the same kind of product, kept its main interaction, and changed where one resource came from. Neither of them listed the kind of product or the subject as something that could change. When I pointed that out, they went back and generated alternatives, and then dismissed whole directions in one line each: one market was crowded, one format needed too much hand-built content, and one was ruled out by quoting a comment I had made about an earlier prototype. I had to ask what production cost even means for a team where the model does all the work.

Asked a third time, they came up with a concept at the right level of abstraction. To show the level I meant, I had listed a few well-known pairings of two ideas in my message. The new concept took its subject from the product that already existed and one half of one of my pairings. The director that proposed it had written, in the same proposal: "The four combinations show a level of abstraction. They are not candidates to imitate." It copied one anyway.

Reordering the steps fixed it. Instructions did not help, which matches what Lou and Sun (2024) found in their anchoring experiments: chain-of-thought, reflection, and explicit instructions to ignore an anchor were each not enough to remove it. So the procedure now runs in four stages.

  1. Goal. Both directors read my words, write what they think I want in parallel without seeing each other, and agree on one goal. The goal names what the user should get and the problems to solve. It does not repeat my examples or the names of anything that already exists.
  2. Proposals. Each director starts in a fresh session and sees only the agreed goal. They read the repository only after they have developed their candidates. Cost, feasibility, and minimum scope are for choosing between candidates, not for deciding which ones get considered.
  3. Discussion. The proposals serve as openings. A conclusion can only be proposed once every question either side raised has been answered.
  4. Check. The director who did not write the conclusion reviews it in a fresh session: does it answer the goal, was anything left unanswered, did anyone change their mind without naming the evidence, were the promised comparisons actually made, and was anything kept from the old product without being compared against an alternative.

I design how the team works in some detail, but I stopped telling the directors what to decide about the product. I send them views instead. One of the things I want to find out with this team is whether product decisions can be left to AI as well, and if I write instructions, the only thing I learn is whether it follows them. So the directors now start each proposal by working out what I am actually after, which of my points are root problems and which are symptoms, and which of my sentences are only examples.

To see whether the reordering worked, I had them redo the decision that had gone wrong, from my original words, without showing them the earlier conclusion. The first attempt leaked: one director listed the board folder, found my request to redo the decision, and read the earlier conclusion in it. A brief now says which single item to read, and nothing else is open until the proposal is written. The second attempt changed the kind of product, the subject, and the main interaction, and kept nothing from what existed. The discussion ran six turns where the earlier one had run two, the director who started out preferring a different concept changed its mind and named the example that changed it, and the check passed.

None of these failures bothered me much. People fail the same way in meetings, and I have been improving LLM workflows for about a year; failing is how they get better. What I liked was the loop. A fix like this used to mean running the development process again. Here I changed a rule beside a running team, the next discussion ran on it, and I could watch whether the fix worked.

Eleven agreements have gone through the check since then. Three came back the first time, each with a specific finding, such as a conclusion that did not answer one clause of what it was supposed to settle, and each passed after the discussion answered it. I also ran one check a second time with the directors' names replaced by "Director A" and "Director B", to see whether a director went easier on a conclusion it had agreed to. It returned the same result.

This is the first shape of it I am satisfied with. Models lean toward whoever is speaking. In a debate, the same answer is judged differently depending on whether it is labeled as the model's own or a peer's, and deferring to the peer is far more common than holding on to its own (Choi, Zhu, and Li, 2025). They also bend their answers toward what the user believes (Sharma et al., 2023). In this team, the words with the most pull are mine. Getting the directors to take my views in without being dragged along by them is, to me, the biggest thing the current shape has done.

One tendency is still there. Left alone, the concepts drift warm: one agreed concept ended its sample scene with an old woman finding the teapot she had shared with her late husband back on its shelf. The things people enjoy draw on unease and transgression as much as on warmth, and the models rarely reach for that end unprompted, so the directors now load a short skill that gives dark and light emotions equal standing when they generate concepts.

How it keeps quality up

The patrol result that surprised me most came from the director-claude reviewer. I still had the orchestrator's session on screen then, and when I looked over between other work, a discussion was running that I had not started. I scrolled back to see why. The product had become a journey through five stations. The reviewer suspected the second half was empty, so it wrote a script that walked every reachable path, ran it for about two minutes, and found that at the last two stations nothing the user did changed anything. It recommended cutting those two stations, and then added: "That item is an agreed plan, so the directors must agree before the route changes." It filed the issue with the script, the command, and the output attached, the orchestrator opened a goal discussion, and the two directors resized the journey. Nobody had asked for any of it.

In every workflow I had built before, an agent that could not decide something escalated to a person, and I wrote those workflows so that escalation stayed rare, a last resort. This reviewer escalated to a discussion between two other models instead. That is one more level of delegation: a decision that used to come to me is now settled between agents. I was glad to see it get this far. I had also not expected it to behave this way, and the flow was so natural that it caught me off guard, in a good way.

That patrol runs every thirty minutes while anything is being built. Studio asks one director to look at what changed since the last one, through two lenses: is anything being added that does not serve the user, or anything the user needs being cut, and is anything being built for cases that will not happen. There have been ninety patrols. Half found nothing. The rest raised something, from real bugs to inconsistencies between the decision records and the code.

It had to be tuned down before it was useful. Its first criteria allowed findings about technical correctness, and it spent its time on those, which ran against its second lens: stop building things just because they are technically correct. The patrol exists to catch cheap models going off course early, not to audit the whole product. Once correctness came out of its criteria, the findings started to be about the product again.

The other thing that keeps quality up is escalation. luna does most of the implementation. When its work falls short, the item goes to sol, a stronger model, which reads why the previous run failed and works from the cause. Seven tasks have gone to sol so far, most of them after luna fell short. Twice luna finished a record of an agreed decision and quietly dropped part of it: once a set of values that an earlier decision had fixed, once an agreed direction deleted during an unrelated edit. In both cases a reviewer caught it, and sol put it back. At first a shortfall went back to luna with more instructions, and a director checked it again afterwards. That was more ceremony than the problem needed, so a shortfall now goes to sol, and sol checks its own work against the task's done conditions.

What is left for me

There are four things I still do. I set the goal of each phase and when it stops. I send views when I have them, in Slack, knowing the directors will weigh them before they decide. I answer questions, and only I can accept a phase as finished. And only I change the team's rules. When an agent's work needs a rule changed, it stops and asks me instead of editing the rules itself. One of them did edit a rule file once, early on. That is why the rule exists.

Most of the time the board is full and my channel is empty. The team sends me a report when a phase needs my acceptance, a question when something only a person can do, and otherwise nothing.

As work, this feels familiar. I work freelance now, but I spent years managing software development, and this is the same job with AI in the engineers' seats. I knew a team like this would become possible.

The speed surprised me. Management 3.0 treats a manager's work as shaping the system a team works in, and that is the work I did here. I had the idea, asked for the system to be built as an experiment, and spent about three days running it and adjusting it. By the end, the team was running its own work, and when a decision went wrong, the changes I made let it catch the next ones itself. Opus 5.5 and GPT-6 Astra can argue a question out and reach a decision once there is a structure around them.

So there is a little unease in this, and a little loneliness. I had expected this future, just not this soon, and it arrived before I had time to get ready for it.

Top comments (0)