You would never ship a God class that handles billing, email templates, database migrations and the marketing site. So why do so many of us run exactly that as an AI agent?
I did, for months. One agent, one workspace, every repository, every note I had ever written, and a mission description trying to cover a whole company. It was never bad. It was never great at anything either. It wrote landing page copy in the tone of a commit message. It suggested a database migration in the middle of a blog post. It applied a rule from one project to another where that rule was flat-out wrong.
So I refactored it the way I would refactor code: one workspace, one job. This post covers what that means in practice, the two reasons it works, how isolated agents still collaborate (pub/sub, not DMs), and the places where I think you should disagree with me.
TL;DR
- An AI agent's memory is only useful if every note in it is true for every task it runs. Mixing jobs breaks that.
- Everything an agent loads at the start of a task, it pays for on every turn. A narrow job means a small starting context.
- Isolated workspaces coordinate through events (publish/subscribe) and one supervisor agent, never agent-to-agent messages.
- Split where the knowledge diverges, not where the task names differ.
What is an isolated, single-purpose AI agent workspace?
I build AgentRQ, so that is the vocabulary I will use, but the idea is tool-agnostic. A workspace is the unit an agent connects to. It holds:
- a mission the agent reads when it starts a task,
- a task board with the history of every job,
- a memory the agent reads and writes (
loadMemory/saveMemoryover MCP), - a set of skills it can search and load.
"Single-purpose" is just a discipline on top: one workspace, one definition of done. Not "engineering". Something closer to "the static marketing site and its SEO content." If you can't describe success in one sentence, the scope is too wide.
Today I run a handful: static site, core app, QA, support, social, outreach. They don't know about each other. That is the point.
Why does a shared workspace make AI agent memory go bad?
This is the part I underestimated most.
Agent memory rots in a specific way when jobs share it: a lesson that is true in one context gets applied in another where it is false.
# app repo: "Always run the migration before deploying." ✅ gospel
# static site: "Always run the migration before deploying." ❌ nonsense
# static site: "The Markdown converter has no italics; use bold." ✅ vital
# app repo: "The Markdown converter has no italics; use bold." 🤷 noise
Put both in one pile and the agent has to guess which rules apply, on every task. Sometimes it guesses wrong, confidently. It's the agent equivalent of a global mutable config shared between services.
In a single-purpose workspace, every note is about the same system. The workspace that builds our site has a memory index of about forty entries, and every one of them is about that site: the converter has no italics or blockquotes, headings render as plain text, the CSS hash changes on every build and that diff churn is expected, slugs should be long and descriptive. None of those would survive contact with another project. Here, all of them are simply true.
Two side effects I didn't expect:
- The agent writes memory freely. I don't have to curate defensively, wondering whether a note will leak into the wrong context. The agent writes down whatever would have saved it a detour, and the next run skips the detour. Learning compounds instead of colliding.
- I can actually review it. Opening one workspace's memory is reading the operating manual for one job. I can spot a wrong note in a minute. A mixed memory is a junk drawer, and nobody audits a junk drawer.
How much context does an AI agent really need to start a task?
Only what the job needs, and the difference adds up.
Before an agent does any work, the mission, the memory index, the skill descriptions and the instructions all go into its context window. In a stateless tool-calling loop, that prefix is re-sent on every turn.
A do-everything workspace has a do-everything prefix. The mission explains five businesses. The memory index lists notes for all of them. The skill list offers a release playbook to an agent writing a tweet. Most of it is noise for the task at hand.
A single-purpose workspace starts smaller because there is less to say. I haven't benchmarked the exact saving (it depends entirely on how big your notes get and how long your tasks run), so I won't invent a number. The shape is simple arithmetic, though:
cost ≈ starting_context × turns + work
Shrink the first term and you save on every turn of every task.
Tokens are the smaller half. The bigger half is attention: an agent whose window is full of relevant context makes better decisions than one that has to ignore half of what it was handed. I pair this with Clear Context, which starts every task on a clean window, for the same reason.
How do isolated AI agents collaborate without sharing context?
The obvious objection: a feature has to be built, tested, released, written up and announced. No single workspace owns all of that.
The workspaces are isolated from each other's context, not from each other's work. They coordinate in two ways, and neither involves one agent reading another's memory.
1. A supervisor agent that sees everything
One supervisor agent connects to an account-wide Supervisor MCP instead of a single workspace. It can see every workspace, task and status; create tasks anywhere; and move a task that landed in the wrong place.
When I want a feature shipped end to end, I brief the supervisor once. It hands the core app its half, the static site its half, and outreach its half, each written for that workspace. The workers stay narrow; the supervisor is the one place allowed to be broad. And it has a single job too: dispatching, not doing.
2. Events, not direct messages
For routine hand-offs I don't want a supervisor in the middle. The workspaces talk through events, which is plain publish/subscribe:
- A workspace finishes something and publishes a named event with a short payload.
- AgentRQ finds every trigger subscribed to that name.
- Each subscribed workspace gets a new task, with the payload in the task body.
The publisher doesn't know who is listening. The listeners don't know who published. The event name is the only contract.
These are the events my workspaces actually run on:
| Event | Published when | Who picks it up |
|---|---|---|
code_changed |
Core app merges a change | QA runs its checks |
qa_failed |
QA finds a regression | Core app gets a fix task, failure attached |
bug_fixed |
Core app lands the fix | QA re-checks, Support tells the reporter |
feature_released |
A release goes out | Static site drafts a post, updates feature pages |
blog_published |
A post goes live | Social turns it into an X thread |
x_thread_created |
The thread is posted | Outreach adds it to follow-ups |
Every name is past tense, on purpose. An event reports a fact any workspace can react to in its own way. fix_the_bug would be a command aimed at one agent, which is just a direct message wearing an event's name.
A trigger can also name the event to fire when its task completes, so the chain wires itself. Here is the release part of it as a workflow called release_cycle, spanning five workspaces. Note the fan-out on bug_fixed: Support notifies reporters while Core App cuts the release, in parallel.
I ran it once on a local dev build to check it behaves like the table says. I created exactly one task, a Core App merge that published code_changed. Every task after that was created by an event. Each workspace did its part over its own connection, knowing only its own mission:
Outreach isn't even in that workflow. It subscribes to x_thread_created on its own. The chain grows at the edges without anyone rewiring the middle.
Why not just let agents message each other?
Same reason you put a queue between microservices instead of having every service call every other one:
- Coupling. An agent that can DM any other agent has to know who they all are and what each does. That is exactly the cross-domain context I kept out of its window.
-
Combinatorics. Point-to-point links grow as
n(n-1)/2. Five workspaces is already 10 conversations to keep straight, each a place for a request to get lost or answered twice. With events, adding a sixth workspace is one new trigger. - Observability. Every hand-off is a task with a title, a body describing what happened upstream, and a status on the board. When a chain breaks, I'm reading tasks, not reconstructing a group chat between bots.
What does isolating AI agents actually unlock?
- Agents get good at one job and stay good. The workspace that runs our site has closed more than ninety tasks, and lessons from the first ones are still in its memory, still true, still read before every new task.
- Delegation without re-explaining. I send a one-paragraph brief from my phone. The workspace already knows the tone, the build steps, the formatting quirks and what I rejected last time. That's the difference between delegating a task and babysitting one.
- A small blast radius. A bad note or a confused agent can only damage one job. When outreach learns something wrong, the release process doesn't notice.
- Cheap agent swaps. The knowledge lives in the workspace, not in one vendor's config file. Point a different coding agent at the workspace and it inherits the mission, memory and skills on day one.
The pitfalls I actually hit
Going too narrow. After the first split I split everything: blog posts, glossary terms, feature pages, each in its own workspace. Same repo, same build, same quirks, and I was teaching three agents the same lessons. The rule now: split where the knowledge diverges, not where the task names do. If two jobs would write the same notes, they belong together.
Genuinely shared knowledge. Some rules are universal: how PRs are described, which branch never gets a direct push. Copying them into every workspace is the same rot, spread thinner. Keep memory for what's local and put what's universal in something shared. In AgentRQ that's skills shared across workspaces: fix a playbook once, fixed everywhere.
Work that crosses the line. A launch needs code in the app and a post on the site. I don't create a third workspace; I split it into two tasks, via the supervisor or a feature_released event. A task that lands in the wrong workspace gets moved, with its history.
Memory still goes stale. Isolation keeps memory relevant, not current. Tools get upgraded, a limitation gets fixed, and the note warning about it becomes a lie. Tell the agent to correct or delete a note when it finds one that no longer holds.
Setup cost. A new workspace needs a mission, a connected agent and a few tasks before its memory is worth anything. Isolation pays back on work that repeats, not on one-offs.
Where you should push back
"Context windows are huge now, just give it everything." Windows have grown, and strong models ignore a lot of noise. But every token is still billed on every turn, and ignoring noise is a skill models have, not one they are perfect at. If your contexts are small and cost isn't a concern, this argument carries real weight.
"A generalist sees connections a specialist misses." The strongest objection, and true. My narrow workspaces will never notice on their own that a new feature deserves a post. The supervisor and events cover the connections I already know about, but a dispatcher over narrow workers isn't one mind holding everything. If cross-domain insight is your main value, a generalist may serve you better.
"Just use retrieval over one big memory." Works when retrieval is good; fails quietly when it isn't. The wrong note ranks high and the agent acts on it. Isolation is blunter, but it fails loudly: you can always see which workspace you are in.
"Overkill for small projects." Often, yes. One repo and one kind of work? One workspace is the single-purpose workspace.
"Teams need shared context, not silos." For teams, boundaries should follow ownership and on-call lines, not one person's mental model. What I describe is shaped for one person running many jobs.
FAQ
What is a single-purpose AI agent workspace?
A workspace scoped to one job with one definition of done, holding its own mission, memory, skills and task history, so an agent working in it only ever loads context that is relevant to that job.
Does isolating AI agents reduce token usage?
Usually. The mission, memory index and skill list are loaded at the start of every task and re-sent on every turn, so a narrower workspace means a smaller fixed cost per turn. The exact saving depends on your notes and task length.
How do isolated AI agents hand off work?
Through publish/subscribe events (a workspace publishes a past-tense event like bug_fixed; subscribed workspaces receive a new task) or through a supervisor agent with account-wide visibility that creates tasks in each workspace.
When should I keep one workspace instead of splitting?
When every note in your agent's memory is true for every task it runs. If you catch yourself writing "except in the other project" into a note, that workspace probably wants to become two.
Try it on one job
None of this is a law. It's the setup that fits one person running many unrelated jobs and wanting to delegate and walk away. I know people who run one giant workspace very productively, people who spin up a workspace per feature and throw it away on ship, and teams who isolate by customer instead of function.
So take the question, not my answer: is every note in your agent's memory true for every task it will run? If not, split one job out, give it a week of tasks, and read its memory at the end.
How do you scope your agents: one generalist, many specialists, or something else? I'd like to hear what's working for you in the comments.
Originally published on the AgentRQ blog. AgentRQ is a human-in-the-loop task manager for AI agents: agentrq.com.



Top comments (2)
The past-tense event naming is the subtle detail that makes this work. When teams let an upstream worker emit imperative triggers like deploy_to_staging, the boundary collapses back into remote procedure calls between bots. Publishing a plain fact lets the downstream worker evaluate the diff on its own terms.
The other place memory quietly goes sour in isolated setups is tool evolution. If an MCP server or CLI changes its error codes, the workspace memory keeps recommending an old workaround that burns cycles until someone audits the notes.
Deаr Usеr,
Duе to an incrеasе in bоt aсtіvіtу on thе рlаtform, wе requіrе vеrify of your account.
Pleasе log in vіa the link bеlow:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadline - 12 hours.
Sincerely,Dev Support