DEV Community

simpny
simpny

Posted on

I Dismissed Multi-Agent Dev Tools as Naive. I Was Judging the Wrong Layer.

A while back I looked at a multi-agent developer tool: a "PM agent", a "dev agent" and a "QA agent" passing messages to each other. I filed it under toys. Agents role-playing an org chart.

I've spent years designing distributed systems. This looked like fan-out with costumes.

I was half right. The half I got wrong is the interesting part.

Surface vs substrate of multi-agent dev tools

The surface often is naive

Agents named after job titles, chatting at each other, burning tokens on pleasantries. Dismissing that was fair. But the persona layer is packaging. I judged the packaging and pattern-matched the whole thing to "toy."

What actually does the work

1. Context isolation. Each agent starts with a clean context window and returns only a summary. The verbose work (searching a codebase, running a test suite, reading logs) never touches the coordinator's window. In an LLM system, attention is the scarce resource. Isolation is how you budget it.

2. Tool scoping. The reviewer can't write files. The tester can't push. Least privilege per role, enforced by configuration, not requested in a prompt.

3. Boundary contracts. The delegation prompt is the request; the summary is the response. Nothing else crosses. That makes it an API design problem, and it's where most of the output quality comes from.

4. Enforced guardrails. A rule in a prompt is a suggestion. A rule in a hook, a script that runs before every tool call and can block it, is policy.

None of these are visible from the demo. All of them are visible from the config files.

Why I missed it

Psychologists call it the Einstellung effect: deep expertise makes the familiar solution so dominant you stop seeing what's new. "It's just orchestration" is true at the architecture layer. Draw the boxes and it's a fan-out.

The constraint lives one layer down, at the operational layer. The workers are non-deterministic and context-bounded. In a microservice, a sloppy interface costs you latency. In an agent system, it costs you correctness, because the downstream agent reasons over whatever you handed it.

Microservices were never interesting as boxes. They were interesting because of failure isolation. Multi-agent systems aren't interesting as personas. They're interesting because of attention isolation.

The rubric I use now

Question Naive signal Real signal
What crosses between agents? Full chat transcripts Typed summaries
How are tools scoped? Every agent has everything Per-role allowlists
Where is governance? In the prompt ("please don't…") Enforced by hooks and permissions
What's the unit of reuse? The persona The agent definition file
How does it fail? Silently wrong output Partial result, flagged and resumable

The takeaway

Evaluate a system at the layer where its constraint lives. For multi-agent tooling, that's not the org chart in the demo. It's what crosses the boundary between agents, and what each one is allowed to touch.

If you've run a multi-agent setup on real work, what broke first?

Top comments (0)