Introduction
I recently attended AWS Community Day Bengaluru on July 11, 2026. There were a number of amazing sessions all through the day but one particular one particularly drew my attention. Himanshu Sangshetti gave a wonderful lecture about designing self orchestrating multi-agent workflows. Creating AI bots is rather easy, scaling them is difficult.
The Problem: One Agent, Too Many Tools
The problem with giving too many tools to a single AI agent? It becomes confused. Give an agent five tools and it will pick the proper one around 95 percent of the time. But scale it up to 30 tools, and accuracy drops to 48%, which is mathematically worse than a coin flip on relevant instruments.
The fix isn’t necessarily a smarter agent. Instead we need some special-purpose agents, each doing one thing exceptionally well, and working together.
Understanding the Protocols: MCP and A2A
The industry is turning to standardized open protocols to tackle this coordination difficulty.
What is MCP?
Model Context Protocol (MCP) In November 2024, Anthropic announced the Model Context Protocol (MCP) as an open standard for how apps might contribute context to huge language models. Consider it the USB-C port for AI. Much like you can plug any device into a USB-C port and expect it to operate, MCP allows models to plug into any tool. It is based on a client-server architecture using JSON-RPC 2.0.
MCP is based on three main primitives:
Tools: Capabilities the model can do, e.g. a search function.
Resources: Files or logs etc. delivered by the program in read-only environment.
Prompts: User controlled reusable templates.
Think of your database as a huge library. MCP is the default library card that allows the AI model to check out what it needs without having to remember the layout of every single building.
What is A2A?
MCP is about how an agent talks to its tools, and A2A (Agent-to-Agent Protocol) is about how agents talk to each other. A2A, donated to the Linux Foundation by Google in 2025, enables opaque agentic programs to communicate.
This enables agents to collaborate without exposing their internal memory, tools or context. It uses an Agent Card (a JSON descriptor at /.well-known/agent.json), utilizes existing standards like HTTP and JSON-RPC 2.0, and enables long running operations that can span from seconds to days.
Imagine A2A like ordering meals on Swiggy. The main app (the organizer) does not know how the dish is prepared in the kitchen by the chef. It basically makes a typical inquiry to the restaurant (the specialty agent), gets status updates (submitted, working, completed) and gets the final result.
Putting Them Together
Add MCP (vertical integration to tools) and A2A (horizontal integration to peer agents) and you have a completely self-orchestrating multi agent system. In this approach the model draws the edges of its own process at runtime.
Building it on AWS Serverless
And now is when things become interesting. How can we construct this architecture on AWS, really?
For the deployment, the agent runtime utilizes Amazon Bedrock with the MCP server and specialty agents hosted on AWS Lambda. A coordinator agent that runs on AWS Step Functions and Bedrock finds the specialty agents and assigns duties over A2A.
You may be wondering if this means writing thousands of lines of complicated code. Thanks to the open source Strands SDK from AWS, it’s remarkably minimal.
Here’s a basic breakdown of the code structure:
# MCP_SERVER.PY
from strands import Agent
from strands.tools.mcp import MCPClient
mcp = MCPClient(
transport="streamable-http",
url="https://...lambda"
)
agent = Agent(
model="bedrock/claude-sonnet",
tools=mcp.list_tools()
)
# COORDINATOR.PY
from strands import Agent
from strands.a2a import A2AClient
a2a = A2AClient()
agents = a2a.discover(
registry_url="https://..."
)
coordinator = Agent(
model="bedrock/claude-sonnet",
sub_agents=agents
)
As you can see, setting up an MCP server or an A2A coordinator requires about 5 lines of code. Plus, if you already have REST APIs, you don’t even need to develop new MCP servers. The AgentCore Gateway can learn your OpenAPI standard from S3 and auto-generate MCP tools in no time.
Four Ways to Orchestrate
In constructing these processes, the team discovered four alternative orchestration solutions using the exact same task, model and tools.
1. Single Agent
Each tool is in an agent on AWS Lambda, invoked directly.
The Good: It is the easiest to develop and debug. No coordination expenses.
The Bad: Tool selection declines badly past about 10 tools.
2. Explicit
Specialist agents work in parallel while Step Functions runs a fixed, hand-authored plan.
The Good: It is quite predictable, which makes it the cheapest and fastest solution .
The Bad: It doesn’t ever let the model plan and can’t adjust if a task unexpectedly changes.
3. Self-Orchestrating
The other agents are the coordinator’s tools and the model routes requests dynamically at run time using A2A.
The Good: It has an infinite potential to flow, and fits perfectly wherever the work is to go.
The Bad: It incurs the highest cost and latency.
4. Hybrid
This defines the self-orchestrating model with a Step Functions guardrail that limits delegations.
- The Good: It maintains runtime flexibility, but limits the "blast radius" with a hard cap (e.g. a maximum of 2 delegations). Guardrails bring you predictability.
The Orchestra-Bench Experiment
To find out which is the most effective technique, the team constructed “Orchestra-Bench”. This is an open benchmark for measuring success, cost, latency across different tactics. They didn't just guess. They ran hundreds of detailed experiments.
That’s what they found over 720 local pilot runs, 53 live AWS runs and 240 public benchmark episodes:
Explicit wins single-shot tasks: On simple jobs, the fixed programmed setting outperformed the other two. Interestingly, as the capabilities of the model drop, the performance difference increases.
Self-Orch wins multi-turn tasks: The winner entirely flipped on complex multi-turn benchmarks. Self-orchestration was the worst option, but now it is the very best with a success rate of 0.60 vs. 0.20 for the clear technique.
The Cascade is Real: In live open-ended jobs on AWS, self-orchestrating agents initiated 6 delegations on average, compared to only 2 with a hybrid guardrail. Total autonomy makes the system around 1.75x slower.
Sharp Edges: What the Docs Don't Tell You
When deploying these architectures to production there are a couple of hidden challenges you need to watch out for:
Cold Start Latency: Multi-hop delegation causes worsening cold starts across several Lambdas. The best fix is to provision concurrency on your critical agents.
Tool Selection Failures: Agents still select the wrong tool 5%-10% of the time. To remedy this, provide highly detailed tool descriptions with few-shot samples.
Context Window Limits: You can’t keep sending all sub-agent outcomes up the chain forever. Protect your memory. Implement defined output schemas and high token budgets.
Debugging Nightmares: When a workflow breaks it can be hard to know which agent failed, and on what input. You will have to rely largely on Step Functions history and AWS X-Ray tracing to debug the problem.
Key Takeaways
So here are the fundamentals to remember while building your own multi-agent workflows, so you don’t get lost in the details:
Task shape dictates your strategy: Always allow the complexity of your mission to determine how you organize your agents.
Read the traces: For numerous hopping agents, good debugging and observability are definitely non-negotiable.
Match the tool to the job: Use self-orchestration when working on large, open-ended problems, and use explicit, scripted routing when working on small, predictable tasks.
Benchmark early: No guessing! Build a simple benchmark using a light-weight harness and assessments to validate your assumptions in the actual world.
Conclusion
In the end, to create scalable multi-agent systems, we need to think beyond the monolithic single-agent systems and toward collaborative, standardized ecosystems. So, the bottom line, if you combine the vertical tool access of MCP and the horizontal delegation of A2A, you can construct really strong AI architectures on AWS that scale well and don’t lose their accuracy.
About the Author
As an AWS Community Builder, I enjoy sharing the things I've learned through my own experiences and events, and I like to help others on their path. If you found this helpful or have any questions, don't hesitate to get in touch! 🚀
🔗 Connect with me on LinkedIn
References
Event: AWS Community Day Bengaluru
Speaker: Himanshu Sangshetti
Topic: A2A + MCP: Building Self - Orchestrating Multi-Agent Workflows on AWS Serverless
Date: July 11, 2026


Top comments (0)