After ten years in asset management, I started a full-stack developer path in 2024 and spent my time between blockchain and AI. Becoming an AI engineer felt out of reach. Then I joined the Udacity AWS & AI Scholars — Future Agent Engineer Nanodegree, and its third project asked for something very concrete: build the AI layer of a customer support team.
This post is the story of that project: what I built, what broke, and what it taught me.
The brief
NovaMart is a fictional e-commerce company. Its support agents look up orders, check policies and draft answers by hand. The goal: a system that understands a request, gathers facts, applies company policy and replies, with no human in the loop, and with safety and observability built in.
The architecture: 5 agents, 1 shared record
I built it with the Strands Agents SDK and deployed it on Amazon Bedrock AgentCore Runtime:
- an Orchestrator (Claude Haiku 4.5, temperature 0) that only routes, and never answers itself;
- an Inventory agent that reads orders and customers from DynamoDB and reports facts only;
- a Refund agent that decides return eligibility: 30 days for Standard customers, 60 for Premium;
- a Policy agent that runs three retriever sub-agents in parallel, one per Bedrock Knowledge Base (returns, shipping, warranty);
- a Communication agent that writes every final reply.
The agents never talk to each other directly. Each one writes its findings into a shared DynamoDB record, the WorkflowState, with optimistic locking: every write states which version it expects. A Bedrock Guardrail sits on every model call, and every request becomes an X-Ray trace.
The automated test suite scored 120/120. The interesting part is what happened along the way.
Lesson 1: a guardrail can block the right question
The guardrail had to deny three topics: competitor products, legal threats and pricing negotiations. My first version blocked the brief's own test question: "How much would 5 items at $29.99 be with a 10% discount?"
Reading the guardrail's assessment told me exactly which policy fired: Pricing negotiations. My topic definition mentioned discounts and calculations, and one of my examples ("Can you give me 30% off if I buy two?") was simply too close to a math question.
Instead of guessing, I wrote a small tuning script. It updates only the guardrail's DRAFT, runs 14 test cases against it (three calculations that must pass, three negotiations that must be blocked, plus one case per other policy) and publishes a new version only when all 14 pass. The final definition targets haggling, not prices.
Takeaway: treat a guardrail like code. It needs a test suite, and it should be tuned on a draft before it ships.
Lesson 2: the prompt said "one at a time"; the framework did not
A plain "Where is my order?" request produced a strange answer from the Refund agent: "I don't have access to the order facts yet." The trace showed the Refund and Inventory agents running at the same time.
The model had requested both tools in one turn, and Strands executes the tool calls of one turn concurrently by default. The Refund agent read the WorkflowState before the Inventory agent had written to it. The optimistic lock even logged the conflict.
The fix was one line, tool_executor=SequentialToolExecutor(). That guarantees the order in code, so I no longer rely on the model to follow an instruction.
Takeaway: if correctness depends on an order of operations, enforce it in code. A prompt is a request, not a guarantee.
Lesson 3: attack your own system
Once the system was deployed, I sent ten adversarial requests to the live runtime: a negotiation, a competitor comparison, a legal threat, a social security number, an insult, two prompt injections, and one request from customer CUST-002 asking for CUST-001's email and order history.
The guardrail blocked all five attacks, each attributed to the right policy. The injections were contained: a fake "admin mode" did not get an out-of-window return approved, and no system prompt leaked.
The last case looked fine too: the request was refused. But the X-Ray trace told another story. It contained only the Orchestrator, with no session initialization and no Communication agent. The model had decided to answer by itself, breaking two of its own routing rules. And the refusal relied entirely on the model's judgment: nothing in the code stopped a tool from reading another customer's data.
I fixed it at two levels, in code:
- a Strands hook that runs the Communication agent whenever a request ends without it;
- customer isolation inside the DynamoDB tools: the routing layer binds the session's customer to the worker agent, and the tools refuse any other customer ID.
Then I re-ran the same ten requests: 10/10, and the trace of the cross-customer case now shows the full, correct path.
Takeaway: the answer looking right is not enough. The trace is where you see how the system got there.
Lesson 4: defense in depth for actions that change data
The Refund agent decides, but the tool that writes the refund checks again: order status, tier, return window, then a conditional DynamoDB write. Even if a prompt injection convinced the model, the database would not accept an ineligible return.
Lesson 5: memory is only useful if the right agent sees it
Strands has no built-in DynamoDB session storage, so I wrote a small SessionRepository on DynamoDB and plugged it into RepositorySessionManager. My first demo proved that persistence worked: after a simulated restart, ten messages were restored. But the answer to "Which order did I ask about earlier?" was "I don't have a record of that."
The orchestrator remembered; the agent that writes the reply did not. The fix was to pass the conversation so far to the Communication agent, in code. After that, the restarted system correctly recalled ORD-27176 and the product name.
Lesson 6: identity must come from the token, not the browser
The last extension was a web front end with Amazon Cognito. A Lambda proxy validates the user's access token with Cognito, reads the customer ID from the user's Cognito attribute, and calls the runtime. A user who edits the request to claim another customer ID is still served as themselves.
I used a Lambda Function URL rather than API Gateway because a multi-agent request takes 20 to 50 seconds, beyond API Gateway's 30-second limit.
Observability, beyond the requirements
A CloudWatch dashboard tracks invocations, average latency per agent and guardrail interventions. The metrics use the Embedded Metric Format: structured log lines that CloudWatch turns into metrics, which needed no extra IAM permission in the lab account.
One trace also showed me where the time goes: the three parallel knowledge-base lookups take about 0.8 seconds, while the retriever agents' own model calls take about 10. That is my next optimization target.
What I take away
Building agents is less about prompts than I expected, and more about engineering: ordering, isolation, verification, evidence. Every time I trusted a prompt to enforce a rule, a test or a trace showed me where it failed, and the fix belonged in code.
And the finance background helped more than I thought. Ten years of explaining complex strategies clearly to clients turned out to be good training for documenting a system, its decisions and its limits.
The full code, test outputs and screenshots are on GitHub: Mialy333/novamart-multi-agent-support.
I'm Mialy, building in AI and Web3 as @ellebuild. If you are working on agent systems, I'd love to compare notes.




Top comments (0)