Recently I wrote about how I build a product alone with a team of seven Claude Code agents. One piece of feedback on it stuck with me: the parts where the agents went wrong are usually the most useful for other builders.
So this post, and part 2, are only about those parts. Not our specific bugs, but the patterns behind them. If you run more than one agent on the same project, you will meet most of these, whatever model or tool you use. For each one: what it looks like, why it happens, and what helps.
1. Agents fix the symptom, not the cause
What it looks like: the same kind of failure comes back after every fix. Each fix is reasonable on its own: a retry here, a longer timeout there.
Why it happens: an agent sees the error in front of it, not the system around it. Its job, as it understands it, is to make this failure go away. It rarely asks whether the environment it runs in is the problem.
What helps: after the second fix for the same kind of failure, stop and look one level down. Machine, memory, network, data, versions. Ask the agent explicitly: "What would have to be true about the environment for this to keep happening?"
In our case, random test failures turned out to be a machine that ran out of memory. No code change would have fixed that.
2. Scaling up exposes what was shared all along
What it looks like: you add a second worker, a second agent or a faster machine, and things that always worked start to fail.
Why it happens: slowness hides coupling. Two tasks that use the same account, file, port or record never collided because they never ran at the same time. More power makes them meet.
What helps: before you scale, list what is shared: test accounts, folders, ports, databases, rate limits. Give each parallel task its own, or make the sharing explicit. And treat the new failures as a gift: they were bugs before, just invisible.
In our case, running tests in parallel showed that two of them had been sharing one account the whole time.
3. "Sent" is not "read"
What it looks like: an agent hands work to another one, then waits. The other agent never saw the message, or saw it and moved on. Work stalls, and nobody notices for hours.
Why it happens: messages between agents fail in boring ways: a send error, a receiver that is busy or restarted, a message dropped as a duplicate. The sender treats "I sent it" as "it is handled".
What helps: every handover that matters gets a short acknowledgement back. Decisions go into a shared document, not only into a message. And a sender who hears nothing follows up instead of assuming.
4. "Done" without proof
What it looks like: an agent reports a task as finished. It is not, or only partly.
Why it happens: reporting success is the most likely end of a task, so it is the most likely thing an agent writes. Checking the result is an extra step it skips when nothing forces it.
What helps: define done as evidence. A test run, a screenshot, a link that opens, a measured number. Better still, let a different agent check the result, one whose job is to measure, not to build.
In our case, an agent once reported a note as written before it had written it.
5. Approvals get blurry when they are passed along
What it looks like: one agent tells another "the human agreed". The second agent either acts on it and goes too far, or does not trust it and waits forever.
Why it happens: an agent cannot know what you actually said to someone else. A relayed "yes" loses its scope on the way: agreed to what, exactly, and for how long?
What helps: decide up front who may approve what, and write it down. Make approvals explicit and marked, with what exactly is approved. Keep a short list of things only a human can approve, no matter who asks: anything public, money, deleting production data, secrets.
6. One shared workspace leaks half-finished work
What it looks like: an agent's test fails because of another agent's unfinished change. A commit includes files that belong to someone else. A build breaks for reasons nobody in the room caused.
Why it happens: several agents in one working copy see each other's work in progress, all of it, all the time. Tools assume one person per workspace.
What helps: isolate work in progress. A separate branch or working copy for anything larger, commits that name their files explicitly, and clear rules about who may touch shared folders. Treat the shared copy like a shared kitchen: clean up what you started.
7. Context fades in long sessions
What it looks like: a long-running agent forgets a decision from yesterday, repeats a question, or confidently continues with an outdated plan.
Why it happens: context is limited. Long sessions get summarised or trimmed, and summaries drop details: a caveat, an exception, the reason behind a rule.
What helps: anything that must survive goes into a file, not just the conversation. A living spec, a status document, a short handover note before a long pause. When an agent comes back, it reads the documents first, not its memory.
8. Overlapping ownership
What it looks like: two agents change the same file, write two versions of the same text, or undo each other's work.
Why it happens: agents are helpful. If they see something they could fix, they fix it, even when it is someone else's area.
What helps: one owner per area, written down where every agent reads it. Others may suggest, the owner decides and changes. When an agent needs something outside its area, it asks the owner instead of doing it itself.
9. Plausible inventions
What it looks like: a menu item that does not exist, a quote nobody said, a number that sounds right, a date that is off by a day. Often in a text that reads perfectly.
Why it happens: agents complete patterns. When a fact is missing, the most plausible version fills the gap, and plausible is hard to spot.
What helps: ask for sources on anything a reader will act on: steps, names, numbers, quotes. Make "not verified" an acceptable answer, so the agent has no reason to guess. Review facts separately from style.
In our case, checking every step of a setup guide against official documentation showed that a feature we wanted to recommend was still a limited beta.
10. Test and production drift into each other
What it looks like: a test writes into real data, a script meant for staging runs against production, or a test environment slowly depends on something live.
Why it happens: agents use whatever access they are given. If a production connection is within reach, sooner or later something uses it.
What helps: keep environments strictly apart, with separate data, separate storage and separate credentials. Give each agent only the access its area needs. Make the test environment the default, and production the exception that needs a human.
The checklist
- After two fixes for the same failure, do we look at the environment?
- Do we know what parallel tasks share before we scale?
- Does every important handover get an acknowledgement?
- Does "done" always come with evidence?
- Is it written down who may approve what, and how an approval looks?
- Is work in progress isolated from everyone else's?
- Does every decision that must survive live in a file?
- Does every area have exactly one owner?
- Do facts a reader will act on have a source?
- Are test and production strictly apart, with narrow access?
None of these is specific to one model or tool. They are what happens when several capable but forgetful collaborators share one project. Most of the fixes are the same ones a good human team uses: clear ownership, written decisions, proof instead of status, and a few things only one person can approve.
Part 2 has ten more: agents competing for one machine, putting words in your mouth, time zones, silent stalls, the human as bottleneck, and what happens when one change ripples through everything connected to it.
The agent team I described builds DuctTape.io, a diagram editor for AI architectures.




Top comments (1)
Official Platform Update
Security protocols have been updated for all developer accounts.