Google's post on the AI Agents Challenge says something it took me years to accept on financial platforms: "multi-agent" was the most frequent claim across thousands of submissions, and a good share of them were a single model walking through a prompt chain with agent names attached. The ones that ranked at the top of each track repeated four engineering decisions that have nothing to do with the model, and everything to do with who operates the system when the primary model returns 503 at 2 AM.
The numbers behind the review
- 40%+: of messages resolved without a model. First tier (regex, zero tokens) of the tiered-routing team, by their own measurement
-
1: validation function for both models.
validate_clinical_response()on the output of Gemini 3.1 Pro and of the Gemini 3.6 Flash fallback -
185 / 24h: default EventBridge retries per target. What a managed bus gives you for free and four
asyncio.Queuein one process do not
What the post is, and what it is not
Sergio Villani's text (Google Cloud AI, September 2, 2026) is not a product announcement, it is a jury report. It describes, without naming teams, the real code of the winners: a telemetry agent that is an MCP client and server at once; a clinical pipeline rebuilt on four asyncio.Queue instances; a clinical-reasoning agent that survived Gemini 3.1 Pro's 503s by falling back to 3.6 Flash without lowering the bar; and a three-tier router where the expensive model only enters after regex and a ten-token classifier have already given up.
The question worth reframing is not "do these patterns work?", they did, there are scored submissions proving it. The question is: what survives when the prototype leaves the hackathon laptop and lands in an AWS account with audit, SLOs and a Bedrock bill at month end? That is the lens of this review. Each pattern below has a version that fits in a notebook and a version that withstands on-call; the post describes the first and hints at the second.
After 16 years operating financial platforms, something else caught my eye: none of the four patterns is new. Tool interface instead of a raw connection, pub/sub instead of a call chain, single-point validation, cheap routing before expensive, this is 2010 distributed-systems engineering with a new vocabulary. What changed is the cost of the mistake: a slow call chain used to delay a screen; now it multiplies tokens.
Pattern 1: bidirectional MCP is an attack-surface decision
The internal half of the pattern is the one worth the most in production. The winning agent did not run SELECT * on the telemetry store and dump rows into context, it reached the database through an MCP tool layer that returned a job's execution plan or a specific stack trace. That is what keeps the context small enough to reason over and what stops a single question from burning the day's token budget. In financial systems I go further: the tool returns an answer bounded by contract (N rows, named fields, no PII columns), because that is what the auditor will read.
The external half is where the post is right and stops early. Exposing the same reasoning as an MCP server turns a chat into infrastructure, another agent, inside an IDE, asks about a job without anyone opening a dashboard. But "that server needs real access control" is a sentence; the MCP specification of 2025-06-18 is a contract: the server is an OAuth 2.1 resource server, publishes metadata per RFC 9728, answers 401 with WWW-Authenticate, and must reject any token whose audience is not itself (RFC 8707). Token passthrough is explicitly forbidden, the server does not forward to the database the token it received from the agent.
On AWS, AgentCore Gateway already delivers both halves: inbound authentication (JWT/OAuth) and outbound (per-tool credential injection), Lambda, OpenAPI and Smithy targets converted into MCP tools, and semantic tool search so the catalog does not blow up the prompt. Use it when the caller is an agent you do not control. If only your own agent calls, a Lambda behind IAM with aws:PrincipalTag does the job, and costs less to maintain.
Pattern 2: four queues in one process are not an event bus
The clinical case is the post's best argument. The linear version, monitoring calls compliance, which calls messaging, which calls dispatch, passed the demo and broke on real use: catching a fall risk from a gait change, cross-referencing a drug-interaction database and reaching someone before the action window closed. In a call chain latency is additive; on a topic bus, two agents that do not depend on each other run at the same instant. A 15% drop in gait velocity publishes CLINICAL.ANOMALY_DETECTED; the compliance agent is already parked on the topic and publishes CLINICAL.COMPLIANCE_REPORT_READY when done, no polling, no explicit handoff.
What the post calls an event bus is four asyncio.Queue instances, one per agent, each with its own coroutine. That is a bus inside one process: the process restarts, in-flight events vanish; a consumer hangs, nobody knows; there is no retry, no DLQ, no ordering guarantee. Fine for the challenge. For production I swap it for EventBridge, the default retry policy per target is 24 hours and up to 185 attempts with exponential backoff and jitter, and the DLQ receives whatever exhausts that, or for SNS fanning out to one SQS queue per agent when I need per-key ordering (FIFO with MessageGroupId = patient).
Two details the in-memory version hides: idempotency, because a managed bus delivers at least once and the messaging agent cannot notify the family twice; and event schema, because a "typed event" in Python is a dataclass, and between services it is a versioned contract in the EventBridge Schema Registry. Whoever forgets the second finds out on the first field change.
The four patterns in a single request
A message comes in, passes through tiered routing, becomes an event, wakes agents in parallel, crosses a single validation function and leaves as an MCP tool for other agents.
🟧 AWS: Roteamento em camadas (padrão 4)
- Camada 0: regex 0 tokens, >40% resolvido (edge)
- Camada 1: classificador ~10 tokens, temperature 0.1 (ai)
- Camada 2: modelo completo só o que sobrou (ai)
📡 AWS: Event bus (padrão 2)
- EventBridge CLINICAL.* por tópico (messaging)
- DLQ (SQS) após 185 retries / 24h (messaging)
🤖 Agentes, reação paralela
- Agente de compliance cruza interação medicamentosa (compute)
- Agente de mensageria idempotente por paciente (compute)
- Agente de despacho (compute)
🔐 Saída, mesma régua (padrões 3 e 1)
- Modelo primário 503 → fallback com backoff (ai)
- Modelo fallback mesma família ou menor (ai)
- validate_response() ponto único, sem atalho (security)
- AgentCore Gateway servidor MCP, JWT + audiência (security)
Flows
- caller -> regex: message
- regex -> cheap: no match
- cheap -> full: ambiguous
- full -> bus: publishes typed event
- bus -> compliance: ANOMALY_DETECTED
- bus -> messaging: COMPLIANCE_REPORT_READY
- bus -> dispatch: DISPATCH_REQUESTED
- bus -> dlq: delivery failed
- compliance -> primary: inference
- primary -> fallback: 503 / throttling
- primary -> validate: response
- fallback -> validate: response
- validate -> mcpsrv: only what passed
- mcpsrv -> caller: MCP tool for other agents
Pattern 3: the fallback goes through the same function or it is not a fallback
The clinical-reasoning team saw Gemini 3.1 Pro return 503 under real load. The common answer would be a retry loop against the same model; they fell back to Gemini 3.6 Flash with backoff. The detail worth stealing is not the fallback, it is where the validation lives: a single validate_clinical_response() that both paths are forced to call, checking that the answer names a real clinical guideline and not plausible-sounding medical language. It is not "remember to apply the bar twice"; it is making it structurally impossible to apply it once.
I have seen the wrong version of this pattern on a payments platform: validation duplicated per path, one copy updated in a sprint, the other forgotten, and the lower-quality answer leaving through the "rare" path during an incident, exactly when nobody is watching. The hard-won lesson: a fallback without shared validation is a silent SLO downgrade, and silent downgrades are the worst kind in a regulated environment, because the log says 200.
On Bedrock the first line of defense against 503 and ThrottlingException is not even switching models: it is calling through a cross-region inference profile (us., eu., apac.), which routes to another region in the same geography with no additional routing cost, the price is the source region's, and CloudTrail records additionalEventData.inferenceRegion for the audit. Only when that is exhausted do I switch models, and then three rules: the modelId that answered goes to the structured log and to a metric (fallback_ratio with an alarm above 5%); the validation is a function with its own tests, not a prompt snippet; and the fallback never receives tools the primary did not.
Pattern 4: tiered routing is FinOps applied to tokens
The team that looked at its own inference bill found what everyone finds: "where's my order" and "cancel my appointment" going through the same full call as a genuinely ambiguous question. The answer was three tiers, regex for navigational intent (zero tokens), a cheap Gemini call with ten output tokens and temperature 0.1 for the ambiguous ones, and the reasoning model only for the rest. The first tier alone handled more than 40% of messages, by their own measurement, before any model was touched.
It is the same maintenance-cost principle I apply to data platforms: the cheapest processing is the one that does not happen. Tier zero is deterministic, testable with a case table, and auditable, with BACEN or LGPD at the table, being able to say "this answer never went through a model" has value of its own. The risk is the inverse: an overly generous regex misclassifies at no cost, and free mistakes do not show up on a cost dashboard. Measure the "regex was right" rate with human sampling, not just the volume.
The managed version on AWS is Bedrock Intelligent Prompt Routing: one endpoint that predicts the response quality of two models from the same family and routes by the responseQualityDifference you configure, with a fallback model as the anchor. It serves as tier 1 or 2, not tier 0, it still spends tokens, and it has three documented limits: it is optimized for English only, it requires exactly two models from one family, and it does not learn from your application's performance data. For the Portuguese in my workloads, tier zero remains regex and tier one remains my own classifier, with a maxTokens of ten and the intent list versioned in the repository.
Where the four patterns shine
- Context bounded by construction: the tool returns the execution plan, not the table, the token budget stops depending on the discipline of whoever writes prompts.
- Latency stops being additive: compliance and messaging wake on the same event; the fastest agent no longer waits for the slowest.
- Fallback without lowering the bar: one validation function, two models, no shortcut, the incident does not become an invisible quality drop.
- Cost cut before inference: more than 40% of traffic resolved by regex, with zero tokens and an auditable answer.
- Reasoning becomes infrastructure: an MCP server is called by other agents with no second integration, as long as it validates token audience.
Where the challenge prototype does not survive: Three things break in the first week of production. In-memory bus: four
asyncio.Queueinstances die with the process, with no retry, no DLQ and no queue metric, swap for EventBridge or SNS/SQS before turning on real traffic. MCP server without audience: accepting any valid bearer token is the specification's confused deputy; validateaudper RFC 8707 and never forward the received token to the database. Regex with no error measurement: tier zero costs no tokens, so its mistakes show up on no bill, sample and count.
How to adopt on AWS, in the order that reduces risk
Measure tier zero before anything else: Export a week of messages, run the regex list offline and count how many match and how many match wrongly. Without that number, pattern 4 is an opinion.
Put the database behind tools with bounded answers: One Lambda per tool, read-only IAM per table, max N rows and named fields. That is the internal half of pattern 1 and the prerequisite for the external one.
Replace the call chain with EventBridge plus DLQ: One topic per typed event, one rule per agent, an SQS DLQ and an idempotency key in
detail. Keep the default retry (24h/185) until you have a reason to shorten it.Write validation as a tested function, then enable the fallback: Cross-region inference profile against throttling first; only then an alternate model. The
modelIdthat answered goes to log and metric with an alarm.Expose the MCP server last, through AgentCore Gateway: Inbound JWT with validated audience, outbound with per-tool credentials, and a single external caller in the first week. A new surface deserves small traffic.
Anti-patterns the post hints at and I have seen up close
- Badge multi-agent: one model, five prompts with agent names, summed latency and no parallelism, a single-threaded system with a new label.
- Per-path validation: one copy for the primary, another for the fallback; the second ages and the incident leaves through the least-watched door.
-
Raw connection in context:
SELECT *dumped into the prompt; works on the demo dataset and blows the token budget on the first real table. - Token passthrough: the MCP server uses on the database the same token it received from the agent, the downstream now trusts a token it never validated.
Curator's note: If I had to pick one of the four to ship tomorrow on a financial platform, it would be pattern 3, the single validation, because it is the only one that protects against a mistake the log does not show. The other three save money or latency; this one saves an incident with an auditor. I would write the validation function first, with tests running in CI against recorded answers from both models, and only then enable the fallback. The lesson I carry from years on call: the rare path is the one that runs on the worst day, and code that runs on the worst day needs the same tests as the happy path.
Verdict
The post is worth keeping as a checklist, not as a reference architecture: it correctly describes what separates multi-agent from prompt chain and stops before the expensive part, which is operating it. Adopt the four patterns when your system has agents on different tempos, a model that has already returned 503 in production and an inference bill someone reads every month. Stay with the call chain when it is two steps, one model and no SLO, the bus and the single validation cost maintenance that case does not pay for. On AWS, the translation is EventBridge with a DLQ, a cross-region inference profile before switching models, and AgentCore Gateway only when the caller is an agent you do not control.
Rating: 4/5
References
- Google Developers Blog: 4 engineering patterns behind the strongest AI Agents Challenge submissions (2 set 2026)
- MCP Specification 2025-06-18: Authorization (OAuth 2.1, RFC 9728, RFC 8707, token passthrough)
- Amazon Bedrock AgentCore Gateway, inbound/outbound auth, alvos Lambda/OpenAPI/Smithy, busca semântica de tools
- Amazon EventBridge: How EventBridge retries delivering events (24h / 185 tentativas)
- Amazon Bedrock: Cross-Region inference (perfis geográficos e globais, sem custo de roteamento)
- Amazon Bedrock: Understanding intelligent prompt routing (responseQualityDifference, limites)
- AWS Machine Learning Blog: Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime (18 set 2026)
Originally published at fernando.moretes.com. By Fernando F. Azevedo: Senior Solutions Architect.
Top comments (0)