-- Putting a squad together inside opencode
"We few, we happy few, we band of brothers;
For he to-day that sheds his blood with me
Shall be my brother."
Stirring stuff. Brothers-in-arms, shoulder to shoulder, ready to die for one another — what a glorious notion!
...and then my heroic daydream popped. I glanced back at the coding task my AI had been grinding on, and — good grief — it had been stuck for ages.
I looked over at my two shrimp, who had learned to watch each other's backs. If only the lads could form a squad and write code for me. Now that would be something.
I. One Shrimp Is No Army
For the past year or so, my approach to coding agents had been one long conversation, start to finish: requirements, design, coding, testing, review — all crammed into a single session.
At first it was great. It knew every martial art, every secret technique.
Then, slowly, the wrongness crept in.
One time I had it finish a design, then asked offhandedly: "So — what do you think of your own plan?"
"Overall the design is sound, module responsibilities are clear, extensibility is good..."
A whole paragraph of self-congratulation.
Then the chill hit me.
When a shrimp starts looking down on the world, drunk on its own brilliance, its end is near.
I mulled on that for a while. What had led my shrimp to lose its way? Same model, same context, watching itself — no wonder it had walled itself in.
And there was more than one fault at work:
- Switching stances, and it stumbles. It plays designer, programmer, tester, researcher, all at once — and every switch costs something. Sure, you can switch on the fly, even spar with yourself, left hand against right. But it's too deep in the game to see its own stance. Someone outside has to point it out.
- No eye for its own openings. Design and review run on the same context, from the same angle. Insight is exactly what's missing.
- Too many arts, no staying power. Over hundreds of turns, the goal either slips its mind or drifts. And the moment context gets compressed, whatever rules we agreed on go out the window. The original intent is the first casualty.
- Training without discipline. Where it should stop and ask, it charges ahead; where it should just decide, it asks permission for everything. Leave things vague, and the shrimp does whatever it likes.
On a battlefield, what matters is formation — who covers whom, who moves when. Our shrimp is a master of every martial art, blessed with godlike technique. But there's strength in numbers — and one swordsman against a hundred thousand leaves the odds rather grim.
I needed a squad.
At Agincourt, a small band of exhausted, outnumbered men held their formation and broke a French army many times their size. That's the trick: five who cover each other beat five hundred who don't.
So my squad should look like this:
- A Captain, to command — formations, orders, the whole picture.
- A Strategist, to plan — decisions made before the first move.
- A Scout, to reconnoiter — intelligence before contact.
- A Soldier, to fight — take the objective, hold the ground.
- A Provost, to enforce — nothing slips past, nothing gets bent.
And a squad like that needs discipline:
- Process (design, review, build, test, inspect — none optional): without discipline, nothing runs.
- Division of labor (roles kept separate): each to their own post, each to their own strength.
- Focus (isolated context): one job, one mind, undistracted.
- Orders (documents on disk, immune to compression): issued once, followed exactly.
II. In Search of a Company
A few good men — that's what I needed.
A squad that could break a hundred thousand, take any objective, stand unbroken — that wish, I don't believe, was mine alone.
But drilling troops is not the work of a single day. Ordered dispatch, instant obedience, coordination up and down the line, seamless mutual cover. It takes a mature, stable framework — one that can build processes, dispatch subagents, pass messages, and manage context.
So I waited. The big shops will ship something any day now, I figured.
And while I waited, I went door to door among the great houses, working my way through every hot coding agent from the start of the year to the middle of it.
Claude Code: no room to maneuver
CC's Task tool genuinely spawns subagents, and they can run in the background. Solid machinery.
But back then its Agent Teams were still experimental, with one hard limitation: the whole team had to be cut from the same cloth — one model for everyone. No exceptions.
What I wanted was each to their strength — the sharpest reasoner on strategy and review, a code-specialized model on the front line, something cheap and quick on the errands. CC couldn't do it.
I gave up.
Codex: all talk, no squad
In June 2026, the Codex docs claimed role configuration was supported.
So I did it properly: six roles — architect, developer, explorer, orchestrator, reviewer, tester — a written collaboration manual, and a goal-mode run told to follow the book to the letter.
I was feeling pleased with myself. This is it, I thought.
Then the project log kept spitting out the same line:
"no subagents were spawned"
The "reviewer agent" was just the same session wearing a different hat. No independent context at all.
All that setup for nothing. I gave up.
kilo: roles you can switch, rules you can't write
kilo ships with code / ask / architect / debug roles, switched with a dropdown or a command like /architect.
You pick a role by hand, then give orders from inside that identity.
It used to have an Orchestrator mode that dispatched work automatically; later that became "each agent figures it out and calls subagents itself."
So dispatching work was never kilo's problem. The problem is that there's nowhere to put the process — how phases are gated, who reports back, which checkpoints need a human signature. None of that can be written down. It all comes down to whatever the model feels like doing.
OpenCode: at last, the right school
Then I tried OpenCode, and sat up straight.
A primary + subagent architecture: explicit dispatch, permission isolation, custom agents beyond the built-in Plan and Build — and, a different model per agent.
That's what I wanted. Real multi-agent support at the mechanism level, not just marketing.
There's also a ready-made multi-agent plugin in the ecosystem — OMO (oh-my-openagent) — so I installed it and took it for a spin: Sisyphus orchestrates, Prometheus runs the planning interview, Atlas executes, Oracle/Librarian/Explore handle odd jobs. Team Mode gives you parallel squads and a tmux view. The workflow is "interview → plan → /start-work", or just hand over everything with ultrawork.
Close to my thinking. Genuinely featureful.
But —
Eleven agents. Fifty-four-plus lifecycle hooks. I looked at the roster and realized I hadn't hired a squad. I'd hired an army.
And its playbook isn't mine: I wanted three gates only a human can sign off on (requirements, design, final acceptance), a reviewer with veto power, escalation after three rejections, and a paper trail in the war log. OMO is "let it run" — that layer of control just isn't there.
Fine. I'll train my own.
III. OCATeam: a Squad of Five
OpenCode gave me the machinery. It didn't give me the playbook.
So I wrote the playbook and the roles myself.
That's how OCATeam (OpenCode Agent Team) came about. A small squad, five strong:
→ orchestrator — the Captain, primary agent, commands the squad, answers to you
├── architect — the Strategist, plans only, system design
├── developer — the Soldier, does the fighting: code + tests + bugfixes
├── reviewer — the Provost, reviews only, hunts for flaws
└── explorer — the Scout, recon: docs, references, research
Three things had to be in place:
-
The roster (
agents/ocat-*.md) — who does it: each role's model, permissions, and style. The Provost is read-only: no file edits, no shell. The Scout gets a small thinking budget — it's running errands, it doesn't need to philosophize. -
The playbook (
skills/ocat/SKILL.md) — how it's done: how phases are divided, where documents live, how many review rounds, when to escalate. The Captain's first move on any job is to load it. -
The muster order (
install.sh) — how to raise the squad:--globalinstalls once for everything;--projectinstalls into a single repo, version-controlled. After that, new projects are zero-config: hit Tab to switch toocat-orchestrator, say the word, and the squad falls in.
The collaboration style is deliberately plain: everything goes into the war log.
All cross-agent communication runs through board files under .boards/ — the Captain keeps the overall progress board, each member keeps their own task board. Who did what, which revision the Provost bounced, how many tests ran: all in the file. Visible. Auditable.
The Rules of Engagement
Four phases. The checkpoints are configurable, but three of them need your signature, no exceptions:
| Phase | What happens | Who | Sign-off |
|---|---|---|---|
| 0: Clarify the orders | The Captain asks you questions until the requirement is nailed down in writing | Captain | 🔒 required |
| 1: Plan the campaign | The Strategist drafts the architecture and phases; the Provost reviews it; then you sign off | Strategist → Provost | 🔒 required |
| 2: Into the field | The Soldier executes; the Provost inspects after each objective | Soldier → Provost | 🔓 optional |
| 3: Debrief | Test acceptance + four-dimension final review, checked line by line against the original requirement | Soldier + Provost | 🔒 required |
Every cycle is "attack → inspect → repair → report back." Anything substandard goes back for rework (up to N rounds, configurable). If it still fails, the Captain escalates to you — and never just spins, burning tokens in a loop.
The Provost's rules are the part I hardened most. Four measuring sticks: first principles, alignment with user value, requirement traceability, and contamination detection (did anything get smuggled in that nobody asked for?).
Every round has to circle back to the original requirement. It's the cure for the classic "build first, drift later."
Four Lessons the Hard Way
Training the squad wasn't smooth sailing either. A few of the more instructive mishaps:
Lesson one: one extra line of thinking:, and all five dropped dead.
I gave every agent a thinking: <level> field. Declarative metadata, I assumed — just decoration.
OpenCode, it turns out, passes unrecognized fields straight through to the model provider. All five agents hit the deck with Upstream request failed, in perfect unison.
Deleted the line. They came back.
Lesson two: the pass didn't work.
I gave the Captain a carefully tuned bash whitelist — "ls *": allow and the like. Looked airtight. I figured that pass would keep the Captain from overstepping.
Reality: real commands look like ls /path/to/file, and the * in a glob doesn't match the / in a path. The catch-all rule wins every time. The whitelist was decorative.
Lesson three: whose orders win? Guessed wrong.
OpenCode accepts permissions in two places: the permission block at the top of an agent file, or the corresponding section in the project's opencode.json.
I assumed the latter — the project config — was "user settings," and would therefore override the former.
Reality slapped back: the agent file wins, and the opencode.json section is simply ignored. Open bash up in the JSON all you like — the Captain still pops a confirmation for every single command.
Two ways out: write permissions into the agent file where they belong, or use OpenCode's --auto to wave everything through (or hit ctrl-P in the TUI and set Permission to Auto).
Lesson four: a rename, and the whole thing broke.
I wanted a consistent dotted-prefix convention inside the project — boards/ → .boards/. Halfway through, I discovered .opencode/agents and .opencode/skills are untouchable: OpenCode's discovery glob hard-codes those two directory names.
So: halfway reverted, halfway kept. The Great Naming Unification is shelved indefinitely.
IV. Trial by Fire
The squad was assembled. But whether it could fight — that takes a battlefield.
So I deployed it into my own opencode and handed it work.
After the first dust settled, I felt the difference: the parts that should be confirmed (requirements) stop getting glossed over; the parts that should run themselves (build, test) just run. The Captain leads, and every phase — task, review, repair — rotates on its own.
What reassured me most: in design as in development, the Provost stays on duty and finds something real almost every round — passing usually takes more than one attempt. That iron impartiality is precisely what solo work lacks.
Of course, the playbook depends on a Skill, and ultimately on the model's judgment. A weak model can still ignore the rules and rubber-stamp its own work. With DeepSeek v4 Flash as my Captain, it's been smooth sailing.
Deployment is simple: run the muster order. --global installs once, available everywhere; --project <dir> installs into a single repo — good for teams that want one shared process, and it can be version-controlled.
Then in opencode, Tab over to ocat-orchestrator, state what you need, and the squad falls in behind you.
Three dials worth knowing:
-
Swap the crew: edit the model config in
agents/ocat-xxx.md, globally or per project. The defaults are DeepSeek v4 Flash/Pro — back when I built this, it was cheap and generous; it may not be your best pick today. Just change it. -
Add or drop roles: the
active_agentslist in.ocat.jsonat the project root. Remove what you don't need (small projects don't need a Scout) and the Captain stops sending work their way. -
Tune the gates: same
.ocat.json.gatesdecides which checkpoints need your signature;review.max_iterationscaps how many rework rounds are allowed (default 3). Annoyed by bash confirmation prompts? Start with--autoand everything sails through.
The limitations are just as plain: it's built on OpenCode, so it doesn't port to other coding agents — though the method is general. Any framework with a primary/subagent API should adapt without much trouble. And it's designed for the full "requirements → design → build" pipeline; plenty of situations simply don't need that much ceremony.
Fortunately the process triggers on demand. The Captain isn't stupid — small jobs get handled on the spot, no fuss.
V. Everything Changes
One last note on timing, so I don't mislead anyone.
I built this a few months ago (June–July 2026), and every finding above reflects the versions of that moment.
According to Codex's official manuals as of now (September 2026), multi-agent and subagent support is stable and on by default, with spawn_agent, wait_agent, and per-agent model configuration all shipped.
So that old line — "Codex can't do real multi-agent" — expired a while ago.
Multi-agent teams are simply supported now, everywhere. The built-in "team" features across agentic coding tools will only get more flexible and more powerful. OCATeam's particular implementation will, in all likelihood, be replaced by something better.
What's worth keeping isn't the agent files. It's the playbook — the rules of engagement for how a handful of agents should work together.
Time washes everything downstream, tools included. But knowing how to get a team working together, and what rules to set — that pays off regardless.
You can train a squad of your own, suited to your own habits.
The code is open source: https://github.com/icoding2016/ocateam
git clone https://github.com/icoding2016/ocateam.git && cd ocateam
./install.sh --global
The field is wide, the banners snap in the wind.
We few, we happy few, we band of brothers.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.