I have been using a simple data bus to let agents working on projects communicate.
It does the obvious things: an agent sends a message, another receives it, and work continues. But with new agent frameworks appearing constantly, I started wondering whether this approach was already outdated.
That question led to a more useful one:
What does my system need beyond delivering messages?
Imagine two agents collaborating on a feature. One implements it. Another reviews it. Then the implementation agent restarts, a message gets delivered twice, and the reviewer approves changes against an old commit.
The agents can still communicate. The project is now in trouble.
This is where the framework discussion becomes interesting.
Four layers that are easy to mix up
Before picking a framework, separate the responsibilities:
| Layer | Question it answers | Examples |
|---|---|---|
| Transport | How does a message reach a worker? | Your service, a queue, HTTP, a message broker |
| Agent interoperability | How does an independent agent discover and request work from another? | A2A |
| Orchestration | Who acts next, and what happens after failure? | LangGraph, application code, workflow engines |
| Tools and context | How does an agent access external capabilities and data? | MCP, ordinary APIs, tool functions |
A bus can deliver a task without understanding its lifecycle. A protocol can describe a task without deciding your project's dependencies. An orchestrator can coordinate workers while using your existing bus underneath.
These layers can coexist.
MCP and A2A: where each fits
MCP and A2A have complementary roles.
MCP equips an agent with tools and context: query a database, inspect a repository, or call a service.
A2A provides a contract for communicating with independent agents: discover capabilities, exchange messages, follow tasks, and retrieve outputs.
For example:
- A developer uses an MCP tool to inspect an issue.
- It delegates an investigation to a remote specialist through A2A.
- Your coordinator records the task and waits for the result.
- A queue buffers execution while the specialist is unavailable.
You can also expose an agent as an MCP tool. That can be perfectly adequate for a bounded request. A2A becomes more relevant when you want the independent agent's task lifecycle and capabilities represented through a common protocol.
Neither protocol automatically settles who may merge code, retry a deployment, or approve a project milestone. Those are application decisions.
What is actually new in 2026?
There are concrete developments worth watching, beyond another framework launch.
A2A now has a stable baseline and a CLI
The A2A project announced its first stable v1.0 specification in March 2026.
On October 1, it introduced an official command-line client. That is useful for scripts, CI jobs, and coding assistants that can run commands:
# Read an agent's capabilities.
a2a card get https://agent.example.com
# Submit a request and receive updates.
a2a send -a https://agent.example.com --stream \
"Review the task contract and identify missing acceptance criteria."
The URL above is illustrative; these commands require an installed CLI and a reachable A2A server, with authentication configured where required. I have not executed them against a live agent for this article.
a2aproject
/
a2a-cli
The official command-line interface for interacting with A2A (Agent2Agent) compatible agents.
A2A CLI
The A2A CLI (a2a) is the official command-line client for A2A (Agent2Agent) agents, maintained by the A2A Project Team.
Why the A2A CLI
A2A is a bidirectional protocol: any A2A agent can act as a client to another and hand off a task. But an agent without an A2A layer can't easily send messages or receive updates. The A2A CLI fills that gap.
With it, you — or any model, agent, or coding harness that can call tools — can discover A2A agent's capabilities and send it work.
Three ways to get started:
-
Empower AI coding assistants — give your coding agent the
a2atool to offload work to remote A2A agents. - Interact instantly — fetch agent cards, send messages, and stream real-time updates from the terminal.
- Automate workflows — seamlessly integrate A2A agents with…
Durable execution is a prominent framework capability
LangGraph documents persistent state and checkpoints. Microsoft Agent Framework documents checkpoints and resuming.
My reading of these capabilities: framework selection increasingly deserves a recovery test alongside the usual delegation demo.
Can the run continue after interruption? Which operations will replay? What state survives? What must your tools make idempotent?
Persistent conversation history helps, but it is only one part of recovering work.
AutoGen advice needs a date attached
The AutoGen repository now identifies the project as being in maintenance mode and directs new projects toward Microsoft Agent Framework.
Existing AutoGen applications may still be useful. For a new build, I would evaluate the successor before following an older tutorial.
Which framework would I choose?
This is a practical shortlist, not a market-share ranking. The fit assessments are my engineering judgment; the links describe the underlying capabilities.
| Option | Documented approach | Where I would evaluate it first |
|---|---|---|
| LangGraph | Stateful graphs, subgraphs, checkpoint storage, multi-agent patterns | Explicit project workflows with recovery requirements |
| Pydantic AI | Tool-based delegation, programmatic handoffs, typed graph control flow | Python services with validated task inputs and results |
| Microsoft Agent Framework | Sequential, concurrent, handoff, group chat, manager-led orchestration | Applications needing several coordination patterns |
| CrewAI | Collaborative Crews and event-driven Flows | Prototyping teams with distinct responsibilities |
| Agno | Coordinate, route, broadcast, and task-based team modes | Specialist teams with a coordinating leader |
| Mastra | TypeScript agents and workflows with suspend/resume | Agent functionality embedded in TypeScript applications |
For my bus-based setup, LangGraph and Pydantic AI are particularly interesting.
LangGraph gives me a place to express transitions: assign work, await implementation, validate results, request review, and finish or retry.
Pydantic AI is attractive when I want ordinary Python functions and validated contracts to remain central. Its documentation distinguishes delegation from handing control to another agent, which is a distinction worth preserving in your own design.
CrewAI and Agno make it easier to express a team. I would still test whether their execution model fits long-running work, independent processes, and project ownership before adopting them.
Ready-to-run agents are another category
An agent framework is something you build with. An agent runtime is something you can run as a worker.
Hermes's current repository includes an A2A integration, with peer discovery, calls, and capability-based orchestration. Availability depends on your installed version; documentation on main is not a guarantee about an older release.
OpenCode supports primary agents and specialized subagents, with configurable prompts, models, and permissions. That makes it a candidate coding worker. Its local delegation features alone do not give your distributed system task ownership or recovery.
Connecting existing workers may be a smaller change than rebuilding them inside a single framework.
Give agents a task contract
The first improvement I would make to a simple bus is to replace ambiguous requests with structured assignments.
Instead of:
Help with the backend.
Send a bounded task with enough information to produce a verifiable result.
This is an illustrative application schema, not an A2A wire-format example. Validate it at your boundary and map it to the protocol or framework you use.Example application task envelope
{
"schema_version": 1,
"message_id": "msg-1042",
"task_id": "task-87",
"project_id": "project-alpha",
"event_type": "task.requested",
"target_capability": "backend.implementation",
"idempotency_key": "task-87-attempt-1",
"deadline": "2026-10-07T18:00:00Z",
"input": {
"repository": "project-alpha",
"base_commit": "abc1234",
"goal": "Add pagination to the orders endpoint",
"acceptance_criteria": [
"Preserve the existing authorization checks",
"Return a stable cursor for subsequent pages",
"Include integration tests for empty and full pages"
]
},
"expected_output": {
"branch": "string",
"commit": "string",
"test_report": "artifact reference"
}
}
Adding an idempotency field does not itself prevent duplicate side effects. Your storage and execution logic have to enforce the behavior.
Likewise, include acceptance criteria because the worker needs a target it can check. A fluent explanation of completed work is weaker evidence than a commit and test report.
Design for the moment a worker disappears
For project agents, I would use a durable task record with states such as:
| State | Meaning |
|---|---|
| Queued | Available for a worker to claim |
| Running | Claimed under a lease, with progress recorded |
| Awaiting review | Output exists and requires validation |
| Completed | Required checks have accepted the output |
| Failed | The attempt ended unsuccessfully |
| Cancelled | Further execution is no longer requested |
These are application states, not prescribed protocol states.
A worker should claim a task atomically. While working, it renews its lease. If the lease expires, the coordinator decides whether to retry, reconcile partial results, or request intervention.
For external side effects, lease ownership alone may be insufficient. A stale worker can keep running after another worker takes over. Use fencing or destination-level idempotency where the operation requires it.
For code, I would also isolate work in branches or worktrees and bind the review to a specific commit. If the implementation changes afterward, the earlier review should not silently apply to the new version.
Measure collaboration before adding more agents
My suggested evaluation starts with a small set of representative tasks and deliberate failures:
- Deliver the same task twice.
- Restart a worker halfway through execution.
- Return an invalid result schema.
- Let a reviewer receive an outdated commit.
- Exhaust the task's time or token budget.
Track accepted outcomes, total cost, retries, and time to completion. Compare the multi-agent workflow with a single-agent baseline doing the same work.
More agents introduce more handoffs. Each handoff deserves evidence that it improves the result.
The architecture I would build next
I would keep the bus, put task ownership and state in durable storage, and introduce orchestration where the workflow needs it.
Then I would add A2A at the boundary for independently deployed agents, and MCP where workers need external tools.
That gives me an incremental path:
- Define a task contract and validate results.
- Add atomic claims, leases, and retry limits.
- Record artifact references and execution history.
- Introduce a framework for complex transitions and recovery.
- Add interoperability without forcing every worker into the same runtime.
Those answers make an agent system easier to trust, regardless of which framework runs the reasoning loop.
How are you coordinating your agents today: a shared queue, a task board, a framework, or direct calls? What failed first when you added a second worker?
Research snapshot: October 6, 2026. Capability claims link to official documentation or repositories. Recommendations and the proposed task design are my assessment, not benchmark results. Forem Liquid embeds follow the DEV editor guide.

Top comments (0)