I have spent the last few months building multi-agent systems that run unattended — watching
inboxes, scraping pages, drafting content, moving money only when a human approves. Here is
the honest list of what made them useful and what was theater.
Concurrency is not the hard part
People love the demo where five agents "think" in parallel. In practice, five agents reading
the same state and racing to write it is just a distributed-systems problem wearing a
costume. The useful concurrency is boring: independent tasks with disjoint resources.
The pattern that works:
- One writer per resource. A single process owns the database/ledger. Others send it requests. This eliminates 90% of the bugs.
- Idempotent steps. Every step can run twice without harm. Store the step id; skip if seen.
- A queue with a dead-letter lane. Failed tasks go somewhere you will actually inspect, not into a silent retry loop.
If you cannot describe a task as "read-only" or "appends one idempotent record", it does not
belong in parallel.
Roles beat prompts
The biggest quality jump came not from a better model but from splitting one big prompt into
narrow roles:
- Researcher — gathers sources, produces a facts file. Never writes prose.
- Writer — drafts from the facts file. Not allowed to invent facts.
- Adversary — tries to falsify the draft; flags unsupported claims.
- Editor — merges, enforces format.
The adversary role is the one people skip and the one that saves you. A model reviewing its
own work is a rubber stamp. A model told "find the weakest claim and attack it" is
genuinely useful.
Give agents a memory that is not the context window
Context windows are for the current task. Long-term memory belongs in files:
-
facts.md— verified claims with sources. -
decisions.md— append-only log of choices and why. -
lessons.md— mistakes and the rule that prevents them.
Before a new task, the agent reads lessons.md. After a failure, it appends one line. This
single habit did more than any model upgrade.
Guardrails that are not optional
- Confirm before external side effects. Sending an email, posting, paying — all gated on an explicit human "yes". The agent prepares, the human approves.
- Spend ceilings. A hard counter that refuses the next paid call once the budget is hit.
- Secret redaction at every boundary. Logs, error messages, and model inputs. Assume the model will echo whatever you paste.
- Read-only by default. Any mutation is a named command, never an ambient capability.
What was theater
- Agents debating each other in a loop. Fun to watch, expensive, and rarely better than one well-scoped pass plus an adversarial review.
- Custom orchestration frameworks before you need them. A queue table and a few functions beat a graph engine until you have dozens of flows.
- Letting agents browse freely. An agent with an open browser and no allowlist will wander into paywalls and captchas. Scope the domains.
The result
The systems that survive contact with reality are small, heavily constrained, and boring.
They do one narrow thing, log everything, ask before touching the outside world, and get a
little smarter from lessons.md. The clever parallel-thinking demos are the ones that break
first.
I build and write about practical agent systems. If you want an autonomous pipeline that
does not need babysitting, get in touch.
Top comments (0)