A while back I looked at a multi-agent developer tool: a "PM agent", a "dev agent" and a "QA agent" passing messages to each other. I filed it under toys. Agents role-playing an org chart.
I've spent years designing distributed systems. This looked like fan-out with costumes.
I was half right. The half I got wrong is the interesting part.
The surface often is naive
Agents named after job titles, chatting at each other, burning tokens on pleasantries. Dismissing that was fair. But the persona layer is packaging. I judged the packaging and pattern-matched the whole thing to "toy."
What actually does the work
1. Context isolation. Each agent starts with a clean context window and returns only a summary. The verbose work (searching a codebase, running a test suite, reading logs) never touches the coordinator's window. In an LLM system, attention is the scarce resource. Isolation is how you budget it.
2. Tool scoping. The reviewer can't write files. The tester can't push. Least privilege per role, enforced by configuration, not requested in a prompt.
3. Boundary contracts. The delegation prompt is the request; the summary is the response. Nothing else crosses. That makes it an API design problem, and it's where most of the output quality comes from.
4. Enforced guardrails. A rule in a prompt is a suggestion. A rule in a hook, a script that runs before every tool call and can block it, is policy.
None of these are visible from the demo. All of them are visible from the config files.
Why I missed it
Psychologists call it the Einstellung effect: deep expertise makes the familiar solution so dominant you stop seeing what's new. "It's just orchestration" is true at the architecture layer. Draw the boxes and it's a fan-out.
The constraint lives one layer down, at the operational layer. The workers are non-deterministic and context-bounded. In a microservice, a sloppy interface costs you latency. In an agent system, it costs you correctness, because the downstream agent reasons over whatever you handed it.
Microservices were never interesting as boxes. They were interesting because of failure isolation. Multi-agent systems aren't interesting as personas. They're interesting because of attention isolation.
The rubric I use now
| Question | Naive signal | Real signal |
|---|---|---|
| What crosses between agents? | Full chat transcripts | Typed summaries |
| How are tools scoped? | Every agent has everything | Per-role allowlists |
| Where is governance? | In the prompt ("please don't…") | Enforced by hooks and permissions |
| What's the unit of reuse? | The persona | The agent definition file |
| How does it fail? | Silently wrong output | Partial result, flagged and resumable |
The takeaway
Evaluate a system at the layer where its constraint lives. For multi-agent tooling, that's not the org chart in the demo. It's what crosses the boundary between agents, and what each one is allowed to touch.
If you've run a multi-agent setup on real work, what broke first?

Top comments (0)