<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Guilherme Dalla Rosa</title>
    <description>The latest articles on DEV Community by Guilherme Dalla Rosa (@roosterdev).</description>
    <link>https://dev.to/roosterdev</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1225931%2Fbc2e57c2-e582-459d-a14d-3f30d5ca8df4.png</url>
      <title>DEV Community: Guilherme Dalla Rosa</title>
      <link>https://dev.to/roosterdev</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/roosterdev"/>
    <language>en</language>
    <item>
      <title>An agent with hundreds of tools</title>
      <dc:creator>Guilherme Dalla Rosa</dc:creator>
      <pubDate>Tue, 06 Oct 2026 14:38:46 +0000</pubDate>
      <link>https://dev.to/roosterdev/an-agent-with-hundreds-of-tools-148c</link>
      <guid>https://dev.to/roosterdev/an-agent-with-hundreds-of-tools-148c</guid>
      <description>&lt;p&gt;Tools are what turned language models into agents, and the Model Context Protocol (MCP) made plugging them in cheap enough that nobody stops at three. Every system the agent has to touch adds its own set, each with dozens of endpoints, so an agent that started with three tools ends up with far more in front of it than any task needs. The trouble is that every one of those definitions is loaded into the context window up front. A good chunk of the context is gone before the first message arrives. And with so many tools that look alike, the model picks the wrong one more often than you'd expect. Nothing fails outright, but the agent takes longer, costs more per turn and, every so often, does the wrong thing without a single error in the logs.&lt;/p&gt;

&lt;p&gt;Connecting that many tools isn't a wiring problem, it's a context problem. What the model sees and when it sees it decides whether the agent works. That decision belongs to the harness, the loop and the tooling around the model, and it's yours whichever model sits inside it. In the pattern that has settled for this, called progressive disclosure, the model never sees the full catalogue.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where all those tools come from
&lt;/h2&gt;

&lt;p&gt;Most of those tools come from a habit that goes back to the first MCP servers, when the usual way to build one was to map every endpoint of an API to its own tool. That's how a single integration turns into a few dozen tools before anyone has asked whether the agent needs them. Neon's recent video &lt;a href="https://www.youtube.com/watch?v=BqRhBq-_kgE" rel="noopener noreferrer"&gt;MCP Just Got a Whole Lot Better&lt;/a&gt; walks through how that habit came about and what Neon changed on their own server as a result, so it's a good place to start if you want the server-side view.&lt;/p&gt;

&lt;p&gt;The tool definitions are billed on every turn, at a discount if prompt caching is on. The MCP client best practices draw the difference between loading them all and fetching them on demand:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F87ug2ofhdilpabhum11t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F87ug2ofhdilpabhum11t.png" alt="Loading every tool up front against progressive discovery" width="800" height="463"&gt;&lt;/a&gt;&lt;/p&gt;
Source: &lt;a href="https://modelcontextprotocol.io/docs/2026-07-28/develop/clients/client-best-practices" rel="noopener noreferrer"&gt;MCP client best practices&lt;/a&gt;, from the Model Context Protocol documentation.



&lt;p&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;The wrong tool is the part of the bill you can't see, and Anthropic's &lt;a href="https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool" rel="noopener noreferrer"&gt;tool search documentation&lt;/a&gt; says Claude's ability to pick the right tool degrades once you exceed 30 to 50 available tools, a line that two integrations of a few dozen tools each already cross. The post that introduced advanced tool use measured the difference tool search makes on Anthropic's own MCP evaluations:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Without tool search&lt;/th&gt;
&lt;th&gt;With tool search&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Opus 4&lt;/td&gt;
&lt;td&gt;49%&lt;/td&gt;
&lt;td&gt;74%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Opus 4.5&lt;/td&gt;
&lt;td&gt;79.5%&lt;/td&gt;
&lt;td&gt;88.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
Source: &lt;a href="https://www.anthropic.com/engineering/advanced-tool-use" rel="noopener noreferrer"&gt;Introducing advanced tool use on the Claude Developer Platform&lt;/a&gt;, Anthropic.



&lt;p&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;The costs don't stop once the agent starts working, because every tool result lands in the context and the model reads past pages of output on every turn. The model gets less reliable as its input grows, even on simple tasks, which Chroma's &lt;a href="https://www.trychroma.com/research/context-rot" rel="noopener noreferrer"&gt;Context Rot&lt;/a&gt; report measured across 18 models. In an agent, that's the rule you agreed ten turns earlier being forgotten, and I've seen it happen with a much smaller set of tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three tools instead of all of them
&lt;/h2&gt;

&lt;p&gt;Progressive disclosure replaces the catalogue with three meta-tools:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;search&lt;/code&gt; takes a plain description of what the model needs and returns the matching names, each with a one-line description&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;get_definition&lt;/code&gt; returns the full schema of a single tool&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;invoke&lt;/code&gt; calls it and returns the result&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Only the definitions the model asked for ever enter the context.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fua3mr3zuo9vnl1fyiu0d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fua3mr3zuo9vnl1fyiu0d.png" alt="The three layers of progressive disclosure" width="799" height="259"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://modelcontextprotocol.io/docs/2026-07-28/develop/clients/client-best-practices#progressive-tool-discovery" rel="noopener noreferrer"&gt;MCP guide&lt;/a&gt; calls the pattern progressive discovery, a name worth knowing before you search for it, and recommends it for clients that connect to many servers.&lt;/p&gt;

&lt;p&gt;Progressive disclosure only pays off once the tool definitions take a share of the context window worth recovering. Below that, or when every tool is used on every request, plain tool calling is the better fit. The guide gives 1% to 5% of the window as an example of where to put the line, while Anthropic draws its own at 10 tools or 10,000 tokens of definitions.&lt;/p&gt;

&lt;p&gt;Once the model asks for something, somebody has to decide what a match is, and the guide lists four ways to do it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Keyword matching, BM25 or a regex, which is simple and works well when tool names and descriptions are descriptive&lt;/li&gt;
&lt;li&gt;Embeddings over the tool descriptions, which handle synonyms and phrasing the keywords miss&lt;/li&gt;
&lt;li&gt;A small, fast model as a sub-agent that picks the tools for the task, which usually works well but can cost more&lt;/li&gt;
&lt;li&gt;A hybrid, scoring across keyword and embedding rankings, or switching strategy by use case&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The four above are what you build when your provider doesn't ship a tool search of its own, or when you need your own ranking.&lt;/p&gt;

&lt;p&gt;One caveat: a search can come back empty, and nothing in the pattern tells the agent what to do about it. Anthropic's tool search returns an empty list rather than an error when nothing matches, so the model does what models do and tries again with a different query, and then again. The stop has to come from the harness, which in Strands means the &lt;code&gt;limits&lt;/code&gt; you pass with each invocation to cap the number of turns and the total tokens a run can use. What happens when the budget runs out is your decision, whether that's a partial answer, a plain "I can't do that" or a hand-off to a person.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the search can live
&lt;/h2&gt;

&lt;p&gt;These three tools can run on the server, on a gateway in front of every server or on the client that runs the loop, and what changes between them is simply who holds the index.&lt;/p&gt;

&lt;h3&gt;
  
  
  On the server
&lt;/h3&gt;

&lt;p&gt;A server can trade one tool per endpoint for those three. Neon's video calls this a layered tool call and describes how a lot of MCP servers moved to it. AWS built its own AWS API MCP Server this way, and it covers the whole API with two tools, &lt;code&gt;suggest_aws_commands&lt;/code&gt; to search and &lt;code&gt;call_aws&lt;/code&gt; to run the command it found:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastmcp&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;FastMCP&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;typing&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Any&lt;/span&gt;

&lt;span class="n"&gt;server&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastMCP&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;AWS-API-MCP&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nd"&gt;@server.tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;suggest_aws_commands&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;suggest_aws_commands&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Suggest AWS CLI commands based on the provided query.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;get_requests_session&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;session&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;session&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;ENDPOINT_SUGGEST_AWS_COMMANDS&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;query&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="nd"&gt;@server.tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;call_aws&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; 
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_aws&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cli_command&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;CallAWSResponse&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Call AWS with the given CLI command and return the result as a dictionary.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="c1"&gt;# call_aws_helper validates the command and applies the security policy
&lt;/span&gt;    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;call_aws_helper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cli_command&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ctx&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nc"&gt;CallAWSResponse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cli_command&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;cli_command&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# loads the read-only operations index, then server.run()
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;
Adapted from &lt;a href="https://github.com/awslabs/mcp/blob/3e0418ff99c5201bec6df4883729fd9f09b856b9/src/aws-api-mcp-server/awslabs/aws_api_mcp_server/server.py#L78-L287" rel="noopener noreferrer"&gt;AWS API MCP Server&lt;/a&gt;, AWS Labs.





&lt;p&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;The search itself isn't in the code, because &lt;code&gt;suggest_aws_commands&lt;/code&gt; sends the query to a service AWS hosts, which returns the most likely CLI commands with their descriptions and parameters, so the search and the definition arrive in one step. AWS has since marked the server as entering end of development and points users to its managed &lt;a href="https://docs.aws.amazon.com/agent-toolkit/latest/userguide/understanding-mcp-server-tools.html" rel="noopener noreferrer"&gt;AWS MCP Server&lt;/a&gt; instead.&lt;/p&gt;

&lt;p&gt;The limit shows up when there are ten servers built the same way, because the agent is then choosing between ten indexes and ten sets of meta-tools. To be fair, Neon's own counter, which Andre Landgraf's post &lt;a href="https://neon.com/blog/give-your-agent-neon-tools" rel="noopener noreferrer"&gt;Give your agent Neon tools&lt;/a&gt; walks through, is that a tool bundling three calls into one workflow saves the agent from making the same mistakes every time, the way an SDK wraps raw endpoints in convenience methods. Those workflow tools still belong on the server.&lt;/p&gt;

&lt;h3&gt;
  
  
  On the gateway
&lt;/h3&gt;

&lt;p&gt;A gateway sits between the agent and every server, so there is one connection and one index whatever is behind it. Solo Enterprise for agentgateway does it with two meta-tools, &lt;code&gt;get_tool&lt;/code&gt; and &lt;code&gt;invoke_tool&lt;/code&gt;, in the setup that Michael Levan, AI Architect at Solo.io, walks through in his post &lt;a href="https://www.solo.io/blog/mcp-progressive-disclosure" rel="noopener noreferrer"&gt;MCP Progressive Disclosure: Save Tokens, Retrieve Schemas&lt;/a&gt;. On AWS, AgentCore Gateway ships the search as a built-in tool, &lt;code&gt;x_amz_bedrock_agentcore_search&lt;/code&gt;, which has to be &lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/gateway-using-mcp-semantic-search.html" rel="noopener noreferrer"&gt;enabled when the gateway is created&lt;/a&gt; and can't be turned on afterwards. In Strands, AWS ships a plugin for that search in its &lt;code&gt;bedrock-agentcore&lt;/code&gt; package:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;mcp_proxy_for_aws.client&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;aws_iam_streamablehttp_client&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;strands&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Agent&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;strands.tools.mcp&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;MCPClient&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;bedrock_agentcore.gateway.integrations.strands.plugins&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AgentCoreToolSearchPlugin&lt;/span&gt;

&lt;span class="n"&gt;mcp_client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;MCPClient&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;lambda&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;aws_iam_streamablehttp_client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;endpoint&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://&amp;lt;gateway-id&amp;gt;.gateway.bedrock-agentcore.&amp;lt;region&amp;gt;.amazonaws.com/mcp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;aws_region&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;us-east-1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;aws_service&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bedrock-agentcore&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;mcp_client&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;plugins&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nc"&gt;AgentCoreToolSearchPlugin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;mcp_client&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;mcp_client&lt;/span&gt;&lt;span class="p"&gt;)])&lt;/span&gt;
    &lt;span class="nf"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Find me afternoon flights to New York&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

From &lt;a href="https://strandsagents.com/docs/integrations/plugins/agentcore-tool-search/" rel="noopener noreferrer"&gt;Amazon AgentCore Tool Search&lt;/a&gt;, Strands Agents docs.






&lt;p&gt;Before each request, the plugin has the model sum up in one line what the user wants, searches the gateway with that line and loads only the tools that come back, so the model never searches for itself. The price is an extra model call per request and one more hop on every tool call.&lt;/p&gt;

&lt;h3&gt;
  
  
  On the client
&lt;/h3&gt;

&lt;p&gt;The guide itself is written for the client, where the host lists the tools from every server once, keeps the definitions on its side and exposes the search, so the servers stay as they are. Strands 1.57 added a coarser version of the pattern as a vended tool, the MCP router:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;strands&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Agent&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;strands.vended_tools&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;make_mcp_router&lt;/span&gt;

&lt;span class="n"&gt;mcp_router&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;make_mcp_router&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;servers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;orders&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://orders.example.com/mcp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;billing&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://billing.example.com/mcp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;max_connections&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="c1"&gt;# Nothing connects yet. The model sees one tool, mcp_router, whose
# description ends "Permitted server names: 'billing', 'orders'."
&lt;/span&gt;&lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;mcp_router&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="nf"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Which orders shipped late this week?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

Adapted from &lt;a href="https://strandsagents.com/docs/user-guide/sdk/tools/vended-tools/#mcp-router" rel="noopener noreferrer"&gt;MCP Router&lt;/a&gt;, Strands Agents docs.






&lt;p&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;The model connects to a server by name, lists its tools and calls one, all through that same tool, so &lt;code&gt;list_tools&lt;/code&gt; does the job of &lt;code&gt;get_definition&lt;/code&gt; for a whole server and &lt;code&gt;call_tool&lt;/code&gt; is &lt;code&gt;invoke&lt;/code&gt;. The router has no search, which leaves the model picking a server by its name alone and reading every definition on it. A &lt;code&gt;search_tools&lt;/code&gt; tool and deferred loading are on the Strands roadmap, in its &lt;a href="https://github.com/strands-agents/harness-sdk/issues/4545" rel="noopener noreferrer"&gt;tool selection design&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Anthropic's tool search and &lt;a href="https://developers.openai.com/api/docs/guides/tools-tool-search" rel="noopener noreferrer"&gt;OpenAI's&lt;/a&gt; are the native version of the full pattern, where you add the provider's search tool to the &lt;code&gt;tools&lt;/code&gt; array and mark the others &lt;code&gt;defer_loading&lt;/code&gt;. On Bedrock, Anthropic's version works through the InvokeModel API but not Converse, which Strands uses, so there the search on top of the router is still yours to build.&lt;/p&gt;

&lt;p&gt;The catch on the client is the prompt cache, because most providers cache the prompt prefix with the &lt;code&gt;tools&lt;/code&gt; array in it, so adding a definition mid-conversation invalidates it and the miss can cost more than the definitions you saved. Of the guide's two fixes, the router takes the one that sends every call through a single stable tool so the array never changes, and the native versions take the other, appending what the search finds after the cached prefix.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code mode
&lt;/h2&gt;

&lt;p&gt;Search cuts the definitions the model sees, but each call still makes a round trip through the model. In code mode, the model writes a script that chains the calls, the script runs in a sandbox and only what it prints comes back. Kenton Varda and Sunil Pai at Cloudflare named it &lt;a href="https://blog.cloudflare.com/code-mode/" rel="noopener noreferrer"&gt;Code Mode&lt;/a&gt;, and the guide calls it programmatic tool calling. Their colleague Matt Carey's &lt;a href="https://blog.cloudflare.com/code-mode-mcp/" rel="noopener noreferrer"&gt;follow-up&lt;/a&gt; pairs it with search to put more than 2,500 Cloudflare endpoints behind two tools.&lt;/p&gt;

&lt;p&gt;In Strands, the &lt;code&gt;strands-harness&lt;/code&gt; package ships it as a built-in tool, &lt;code&gt;programmatic_tool_caller&lt;/code&gt;, which turns every other tool the agent has into a function the script can call. With the gateway search plugin loading the tools first, the model calls them from one script:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;strands_harness&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;create_harness&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;bedrock_agentcore.gateway.integrations.strands.plugins&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;AgentCoreToolSearchPlugin&lt;/span&gt;

&lt;span class="c1"&gt;# mcp_client: an MCPClient for the AgentCore Gateway
&lt;/span&gt;&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;mcp_client&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;create_harness&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;plugins&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nc"&gt;AgentCoreToolSearchPlugin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;mcp_client&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;mcp_client&lt;/span&gt;&lt;span class="p"&gt;)],&lt;/span&gt;
        &lt;span class="n"&gt;builtin_tools&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;*&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;programmatic_tool_caller&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timeout&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Which team members exceeded their Q3 travel budget?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# The model sends one script, such as:
#   import asyncio, json
#   team = await hr___get_team_members(department="engineering")
#   expenses = await asyncio.gather(*[
#       hr___get_expenses(user_id=m["id"], quarter="Q3") for m in team
#   ])
#   ...  # compare each total with its budget, inside the sandbox
#   print(json.dumps(exceeded))
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

Adapted from &lt;a href="https://strandsagents.com/docs/user-guide/harness/configure/tools-and-instructions/" rel="noopener noreferrer"&gt;Add tools and instructions&lt;/a&gt;, Strands Agents docs, with the model's script from &lt;a href="https://www.anthropic.com/engineering/advanced-tool-use" rel="noopener noreferrer"&gt;Introducing advanced tool use on the Claude Developer Platform&lt;/a&gt;, Anthropic.






&lt;p&gt;&amp;nbsp;&lt;/p&gt;

&lt;p&gt;The script runs in Monty, Pydantic's Python interpreter written in Rust, with no access to the filesystem, the network or the environment, so the credentials stay with the host and the gateway. On the server side, the managed AWS MCP Server offers the same idea as &lt;code&gt;aws___run_script&lt;/code&gt;, where the script reaches only the AWS APIs and IAM checks each call.&lt;/p&gt;

&lt;p&gt;Code mode opens a code execution surface that's now yours to run. The guide is explicit that approving a script doesn't approve every call it makes, so in &lt;code&gt;strands-harness&lt;/code&gt; each call still goes through the agent's hooks and policies, then through the gateway's own authorisation. A call that needs a person's approval fails, though, because the package can't pause a running script to ask for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to put it
&lt;/h2&gt;

&lt;p&gt;In my opinion, the search belongs in the harness you own, whether that's the client or the gateway. Jiquan Ngiam, co-founder and CEO of MintMCP, which sells an MCP gateway, &lt;a href="https://x.com/JiquanNgiam/status/2001461846717661438" rel="noopener noreferrer"&gt;is blunter about the servers&lt;/a&gt;, writing that "progressive tool disclosure should be done by the agent, not as an MCP server tool".&lt;/p&gt;

&lt;p&gt;A small API with a handful of endpoints is still fine with one tool per endpoint. To be fair, a single API as large as Cloudflare's or AWS's is where Matt's case for putting code mode on the server makes sense. As always, it depends on your requirements. If you want to go further, the Strands docs on &lt;a href="https://strandsagents.com/docs/user-guide/harness/tools/programmatic-tool-calling/" rel="noopener noreferrer"&gt;programmatic tool calling&lt;/a&gt; explain how code mode runs in the harness you own.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>mcp</category>
      <category>strands</category>
    </item>
    <item>
      <title>Aurora DSQL now supports foreign keys. What else?</title>
      <dc:creator>Guilherme Dalla Rosa</dc:creator>
      <pubDate>Mon, 07 Sep 2026 10:35:29 +0000</pubDate>
      <link>https://dev.to/roosterdev/aurora-dsql-now-supports-foreign-keys-what-else-2dp0</link>
      <guid>https://dev.to/roosterdev/aurora-dsql-now-supports-foreign-keys-what-else-2dp0</guid>
      <description>&lt;p&gt;AWS recently announced that &lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/08/aurora-dsql-foreign-key-constraints/" rel="noopener noreferrer"&gt;Amazon Aurora DSQL now supports foreign key constraints&lt;/a&gt;. If you've been following DSQL since the preview, you know why this one matters.&lt;/p&gt;

&lt;p&gt;Aurora DSQL is the serverless, distributed database from AWS that speaks PostgreSQL: no instances to size, scaling to zero when idle, and multi-region clusters where you write to any region and read a consistent answer from all of them. Instead of locking rows, it detects conflicts at commit time and asks the loser to retry. I covered all of that in my talk &lt;a href="https://youtu.be/kJG2VEAikuk" rel="noopener noreferrer"&gt;&lt;strong&gt;What DSQL? Rethinking SQL for the Serverless, Distributed Age&lt;/strong&gt;&lt;/a&gt; at AWS Community Summit Manchester, and one of the limitations I highlighted there was the lack of foreign keys, a potential blocker for some use cases.&lt;/p&gt;

&lt;p&gt;The syntax is the one you already know from PostgreSQL: a column-level &lt;code&gt;REFERENCES&lt;/code&gt; or a table-level &lt;code&gt;FOREIGN KEY&lt;/code&gt;, with the same match types, the same five referential actions and the same deferrable options, so the DDL you wrote for PostgreSQL should run as it is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;customers&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;customer_id&lt;/span&gt; &lt;span class="n"&gt;uuid&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;customers&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;ON&lt;/span&gt; &lt;span class="k"&gt;DELETE&lt;/span&gt; &lt;span class="k"&gt;RESTRICT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="nb"&gt;numeric&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;18&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The one DSQL-specific detail is adding a constraint to a table that already has rows. Where PostgreSQL would scan the table there and then, DSQL has you &lt;a href="https://docs.aws.amazon.com/aurora-dsql/latest/userguide/release-notes.html" rel="noopener noreferrer"&gt;add the constraint as &lt;code&gt;NOT VALID&lt;/code&gt; and validate the existing data afterwards&lt;/a&gt;, as an asynchronous job you can follow in the &lt;a href="https://docs.aws.amazon.com/aurora-dsql/latest/userguide/working-with-create-index-async.html" rel="noopener noreferrer"&gt;&lt;code&gt;sys.jobs&lt;/code&gt; system view&lt;/a&gt;, the same way you follow an asynchronous index build:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;ALTER&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;
  &lt;span class="k"&gt;ADD&lt;/span&gt; &lt;span class="k"&gt;CONSTRAINT&lt;/span&gt; &lt;span class="n"&gt;orders_customer_fk&lt;/span&gt;
  &lt;span class="k"&gt;FOREIGN&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;REFERENCES&lt;/span&gt; &lt;span class="n"&gt;customers&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;VALID&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What changes in practice is who enforces the relationship: from now on the database refuses the write, and a schema you carry over from a PostgreSQL project needs fewer exceptions. One caveat: the &lt;a href="https://docs.aws.amazon.com/aurora-dsql/latest/userguide/working-with-foreign-key-constraints.html" rel="noopener noreferrer"&gt;foreign key documentation&lt;/a&gt; is clear that every write to a referenced or referencing table costs extra reads. A &lt;code&gt;CASCADE&lt;/code&gt; also counts toward the 3,000-row limit of a transaction, so a delete that fans out into thousands of child rows still has to be chunked.&lt;/p&gt;

&lt;p&gt;To be fair, not everyone wants foreign key constraints in the first place. PlanetScale's guide to &lt;a href="https://planetscale.com/docs/vitess/operating-without-foreign-key-constraints#why-does-planetscale-not-recommend-constraints-" rel="noopener noreferrer"&gt;operating without foreign key constraints&lt;/a&gt; explains why they don't recommend them. Constraints mean more locking under high concurrency, column types you can no longer change, more complex schema refactors and rules that are hard to maintain once data is split across servers.&lt;/p&gt;

&lt;p&gt;Their advice is to keep the relationships in your model and enforce them in the application, which is what most of us had been doing on DynamoDB anyway. DSQL's implementation avoids the locking part of that list, and the rest turns into the extra reads and the row limit above. Nice to have the option, and still worth measuring before you turn it on everywhere.&lt;/p&gt;

&lt;p&gt;Then, last month at AWS Community Day Singapore, I watched &lt;a href="https://www.linkedin.com/in/yama3133/" rel="noopener noreferrer"&gt;Yuuki Yamashita&lt;/a&gt;'s talk &lt;a href="https://speakerdeck.com/yama3133/distributed-transactions-under-fire-building-a-zero-oversell-flash-sale-platform-with-amazon-aurora-dsql" rel="noopener noreferrer"&gt;&lt;strong&gt;Distributed Transactions Under Fire: Building a Zero-Oversell Flash Sale Platform with Amazon Aurora DSQL&lt;/strong&gt;&lt;/a&gt;. The problem is familiar to anyone in e-commerce: a limited drop lasts 30 seconds, race conditions cause oversells and provisioning for that peak means paying for it all month. So he built a flash sale on Lambda and DSQL where two buyers race for the last item and the first to commit wins.&lt;/p&gt;

&lt;p&gt;His demo put 100 concurrent buyers against 10 items and came out with 10 orders and zero oversells. The honest part of the talk was what broke on the way. Connections cached across Lambda invocations outlived the 15-minute IAM token, and the retries themselves tripled the load until backoff with jitter and a cap of three attempts turned the storm into a clean sold-out. It's the kind of real-world lesson I look for, and it left me wondering what else had improved in the DSQL world since my talk. So here's what I found.&lt;/p&gt;

&lt;h2&gt;
  
  
  What else has improved
&lt;/h2&gt;

&lt;p&gt;Let's start with the one I was hoping for when I first looked at DSQL: &lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/07/amazon-aurora-dsql-cdc-ga/" rel="noopener noreferrer"&gt;change data capture&lt;/a&gt; (CDC). DynamoDB Streams has spoilt me: the database propagates every change to a stream, and the rest of the event-driven architecture hangs off it without the application having to publish anything. With a traditional relational database you end up building the transactional outbox pattern instead, writing the event in the same transaction as the data, polling the outbox table and pushing it to the bus, and keeping the two in step.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F38vvupyyyj6bhl536j7o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F38vvupyyyj6bhl536j7o.png" alt=" " width="800" height="370"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;CDC now does that for you: inserts, updates and deletes go to Kinesis Data Streams as change events, and from there to Lambda, or to S3, Redshift and OpenSearch through Firehose. In my opinion this is as big an announcement as the foreign keys, because it makes the case for DSQL as a replacement for DynamoDB in event-driven applications, with SQL on top. One caveat: delivery is at least once, so the &lt;a href="https://docs.aws.amazon.com/aurora-dsql/latest/userguide/cdc-setup.html" rel="noopener noreferrer"&gt;consumer still has to deduplicate and order the records&lt;/a&gt;, and the stream is billed in DPUs by the volume it captures. If you want to try it, Vijay Karumajji's &lt;a href="https://aws.amazon.com/blogs/database/getting-started-with-change-data-capture-in-amazon-aurora-dsql/" rel="noopener noreferrer"&gt;getting started guide&lt;/a&gt; walks through the setup, from the Kinesis stream and the IAM role to the first events.&lt;/p&gt;

&lt;p&gt;The rest of the engine changes closed several other items on my slide, and the &lt;a href="https://docs.aws.amazon.com/aurora-dsql/latest/userguide/release-notes.html" rel="noopener noreferrer"&gt;release notes&lt;/a&gt; are the place to follow them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Multi-region is no longer a US-only story. &lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/07/amazon-aurora-dsql-adds-multi-region-clusters-four-more-regions/" rel="noopener noreferrer"&gt;Multi-region clusters run in 16 regions&lt;/a&gt; across three region sets, with Frankfurt, Ireland, London, Paris, Spain and Stockholm on the European side, and single-region clusters are available in 20 regions. A cluster still has to stay inside one region set.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/02/amazon-aurora-dsql-adds-identity-columns-sequence/" rel="noopener noreferrer"&gt;Identity columns and sequences&lt;/a&gt; are in, with a cache you have to set explicitly and values that can arrive out of order across sessions.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/05/aurora-dsql-json-support/" rel="noopener noreferrer"&gt;JSON&lt;/a&gt; and &lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/06/amazon-aurora-dsql-supports-jsonb/" rel="noopener noreferrer"&gt;JSONB&lt;/a&gt; are supported, another item from the slide.&lt;/li&gt;
&lt;li&gt;Schema changes got easier: &lt;code&gt;DROP COLUMN&lt;/code&gt;, constraints added as &lt;code&gt;NOT VALID&lt;/code&gt; and validated later, indexes on expressions and extended statistics are all in.&lt;/li&gt;
&lt;li&gt;Clusters now &lt;a href="https://aws.amazon.com/about-aws/whats-new/2025/12/amazon-aurora-dsql-cluster-creation-in-seconds" rel="noopener noreferrer"&gt;create in seconds&lt;/a&gt; instead of minutes, and storage goes up to 256 TiB.
What hasn't moved is the other half of the slide, and the &lt;a href="https://docs.aws.amazon.com/aurora-dsql/latest/userguide/working-with-postgresql-compatibility-migration-guide.html" rel="noopener noreferrer"&gt;migration guide&lt;/a&gt; is the honest place to read it. No triggers, no PL/pgSQL, no temporary tables, no extensions (so no pgvector and no PostGIS), one database per cluster, and 3,000 rows and 5 minutes per transaction. Those aren't gaps waiting to be filled, they're the design, and the examples below are mostly about building around them.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What the community learned
&lt;/h2&gt;

&lt;p&gt;The lesson that repeats across every story I read is that retries are a design decision, not an error handler. Marc Bowes, who works on DSQL, shows in &lt;a href="https://marc-bowes.com/dsql-avoid-hot-keys.html" rel="noopener noreferrer"&gt;avoid hot keys&lt;/a&gt; why a single counter row that every transaction updates is the classic mistake, and why appending rows and summing them is the shape that scales. Fernando Azevedo's &lt;a href="https://dev.to/fernando_azevedo_6844e930/aurora-dsql-multi-region-field-notes-for-financial-grade-systems-2fpj"&gt;field notes on multi-region&lt;/a&gt; add the other rule of thumb: every commit in a multi-region cluster pays the round trip between the regions, so an abort rate creeping up is a design smell before it's a database problem. For the bigger picture, Marc Brooker's &lt;a href="https://brooker.co.za/blog/2025/11/02/thinking-dsql.html" rel="noopener noreferrer"&gt;DSQL: Simplifying Architectures&lt;/a&gt; makes the case for an active-active setup with no failover logic and no leader election. The team also published &lt;a href="https://arxiv.org/abs/2607.13276" rel="noopener noreferrer"&gt;the paper&lt;/a&gt;, for anyone who wants the full story of how it works.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use cases and examples
&lt;/h2&gt;

&lt;p&gt;A few references worth keeping, from talks and production stories to sample apps:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://www.linkedin.com/in/vadymkazulkin/" rel="noopener noreferrer"&gt;Vadym Kazulkin&lt;/a&gt;, AWS Serverless Hero, has been covering DSQL from the Java side on stage and in writing for a while. His seven-part series &lt;a href="https://dev.to/aws-heroes/serverless-applications-on-aws-with-lambda-using-java-25-api-gateway-and-aurora-dsql-part-1-2g27"&gt;Serverless applications on AWS with Lambda using Java 25, API Gateway and Aurora DSQL&lt;/a&gt; goes from the sample application to SnapStart with DSQL request priming and GraalVM Native Image, with the &lt;a href="https://github.com/Vadym79/aws-lambda-java-25" rel="noopener noreferrer"&gt;code on GitHub&lt;/a&gt;, and his re:Invent session &lt;a href="https://dev.to/aws/dev-track-spotlight-build-modern-applications-with-amazon-aurora-dsql-dev308-4g5o"&gt;Build modern applications with Amazon Aurora DSQL&lt;/a&gt; has the latency numbers for an ordering app on single-region and multi-region clusters.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://www.linkedin.com/in/darryl-ruggles/" rel="noopener noreferrer"&gt;Darryl Ruggles&lt;/a&gt;' &lt;a href="https://darryl-ruggles.cloud/dsql-kabob-store" rel="noopener noreferrer"&gt;multi-region Kabob Store&lt;/a&gt;, an e-commerce sample with the Terraform to reproduce it, and his &lt;a href="https://dev.to/aws-builders/amazon-aurora-dsql-a-practical-guide-to-awss-distributed-sql-database-2n58"&gt;practical guide to Aurora DSQL&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;The &lt;a href="https://github.com/aws-samples/aurora-dsql-samples" rel="noopener noreferrer"&gt;aurora-dsql-samples&lt;/a&gt; repository from AWS, with client examples for most languages and ORMs, a booking API with the retry logic in place and a sample AI agent that uses DSQL as its store.&lt;/li&gt;
&lt;li&gt;The AWS Database Blog on DSQL for &lt;a href="https://aws.amazon.com/blogs/database/amazon-aurora-dsql-for-gaming-use-cases/" rel="noopener noreferrer"&gt;gaming&lt;/a&gt;, for &lt;a href="https://aws.amazon.com/blogs/database/amazon-aurora-dsql-for-global-scale-financial-transactions/" rel="noopener noreferrer"&gt;financial transactions&lt;/a&gt; and as the store behind &lt;a href="https://aws.amazon.com/blogs/database/building-an-ai-powered-grid-investigation-agent-with-aurora-dsql-and-amazon-bedrock-agentcore/" rel="noopener noreferrer"&gt;an AI agent on Bedrock AgentCore&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aws</category>
      <category>database</category>
      <category>serverless</category>
    </item>
    <item>
      <title>The Top Serverless Announcements from AWS re:Invent 2023</title>
      <dc:creator>Guilherme Dalla Rosa</dc:creator>
      <pubDate>Mon, 04 Dec 2023 19:15:30 +0000</pubDate>
      <link>https://dev.to/roosterdev/the-top-serverless-announcements-from-aws-reinvent-2023-2d3e</link>
      <guid>https://dev.to/roosterdev/the-top-serverless-announcements-from-aws-reinvent-2023-2d3e</guid>
      <description>&lt;p&gt;With numerous AI-related announcements, this year's re:Invent marked a significant shift towards AI. The standout was &lt;a href="https://aws.amazon.com/about-aws/whats-new/2023/11/aws-amazon-q-preview/" rel="noopener noreferrer"&gt;Amazon Q&lt;/a&gt;, which was revealed during the keynote. Amazon Q, AWS's counterpart to ChatGPT, integrates into the AWS console, AWS documentation pages, and IDEs via the VS Code plugin, AWS Toolkit. Uniquely trained on AWS documentation and immune to the restrictions of &lt;code&gt;ai.txt&lt;/code&gt;, Amazon Q promises more current answers than ChatGPT.&lt;/p&gt;

&lt;p&gt;Steering away from the AI buzz (if that's even possible), let's dive into the serverless realm and check groundbreaking announcements that were made and how they might revolutionise our tech toolkit.&lt;/p&gt;

&lt;h2&gt;
  
  
  ElastiCache "serverless"
&lt;/h2&gt;

&lt;p&gt;The launch of &lt;a href="https://aws.amazon.com/about-aws/whats-new/2023/11/amazon-elasticache-serverless/" rel="noopener noreferrer"&gt;Amazon ElastiCache Serverless&lt;/a&gt; marks a significant stride in AWS's serverless offerings. This new service addresses many limitations of the traditional ElastiCache, making it more user-friendly and fitting the serverless model more closely. &lt;/p&gt;

&lt;p&gt;The key highlights include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Simplified Operation&lt;/strong&gt;: No need to choose instance types or worry about bandwidth and TPS limits.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Native API Support&lt;/strong&gt;: Supports Memcache and Redis APIs, easing migration from server-based setups.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Multi-AZ and VPC Support&lt;/strong&gt;: Offers built-in high availability and works with VPCs from day one.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Eliminated Autoscaling Groups&lt;/strong&gt;: Autoscaling is more straightforward, although it can take time to scale up during sudden spikes.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Pay-per-use Pricing&lt;/strong&gt;: A move towards a dynamic pricing strategy aligned with serverless computing principles.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;While Amazon ElastiCache Serverless introduces several improvements over its predecessors, it has sparked mixed feelings within the tech community, particularly concerning its pricing and its serverless credentials.&lt;/p&gt;

&lt;p&gt;The pricing, notably high, includes a minimum charge of $90 per month for even minimal data storage. For instance, storing just slightly over 1 GB &lt;strong&gt;can cost $180 monthly!&lt;/strong&gt;. This contrasts with on-demand instances, where comparable storage is significantly cheaper.&lt;/p&gt;

&lt;p&gt;Operational concerns also come into play. The requirement to run Lambda functions within a VPC leads to additional VPC-related expenses. Additionally, the scaling capability, which only allows doubling capacity every 10 minutes, is perceived as sluggish. This combination of high costs and operational limitations has sparked debates about the service's practicality and affordability, particularly for sporadic usage scenarios.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lambda Scales 12x Faster
&lt;/h2&gt;

&lt;p&gt;AWS Lambda has &lt;a href="https://aws.amazon.com/blogs/aws/aws-lambda-functions-now-scale-12-times-faster-when-handling-high-volume-requests/" rel="noopener noreferrer"&gt;revolutionised its burst concurrency limits&lt;/a&gt;, significantly enhancing its scaling capabilities. Previously constrained by a region-wide burst limit of 500–3000 and a slow refill rate, Lambda now allows each function to burst to 1000 concurrent executions instantly. What's more, this limit increases by 1000 &lt;strong&gt;every 10 seconds&lt;/strong&gt;, with each function scaling independently.&lt;/p&gt;

&lt;p&gt;This change is a game-changer for scenarios with sudden traffic spikes, like flash sales. For instance, with an average request time of 100ms, a single execution can handle 10 requests per second. So, you can now burst to 10,000 requests per second per endpoint, with an additional 10,000 every 10 seconds, up to your account-level limit.&lt;/p&gt;

&lt;p&gt;This level of scalability introduces new considerations, especially around system bottlenecks. For example, API Gateway, with its default limit of 10,000 requests per second, could now become a throttle point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step Functions Enhancements
&lt;/h2&gt;

&lt;p&gt;AWS Step Functions has introduced several significant updates:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://aws.amazon.com/blogs/aws/external-endpoints-and-testing-of-task-states-now-available-in-aws-step-functions/" rel="noopener noreferrer"&gt;Public HTTP Endpoints&lt;/a&gt;: Step Functions can now directly call any public APIs, eliminating the need for Lambda or API Gateway proxies. They utilise existing HTTP connections from EventBridge.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://aws.amazon.com/blogs/aws/external-endpoints-and-testing-of-task-states-now-available-in-aws-step-functions/" rel="noopener noreferrer"&gt;Testing Individual States&lt;/a&gt;: You can test individual states in your state machine without full execution. This is facilitated by the new TestState endpoint, enabling programmatic testing.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://aws.amazon.com/about-aws/whats-new/2023/11/aws-application-composer-step-functions-workflow-studio/" rel="noopener noreferrer"&gt;Integration with AWS App Composer&lt;/a&gt;: Step Functions now integrates with &lt;a href="https://aws.amazon.com/application-composer/" rel="noopener noreferrer"&gt;AWS App Composer&lt;/a&gt;, allowing for easy inclusion and editing of state machines within stacks.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://aws.amazon.com/about-aws/whats-new/2023/11/aws-step-functions-optimized-integration-bedrock/" rel="noopener noreferrer"&gt;Optimised Integration with Bedrock&lt;/a&gt;: Enhanced support for AI app development, though Lambda remains preferable for streaming responses, especially for frontend applications.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  SQS FIFO Throughput Massive Increase
&lt;/h2&gt;

&lt;p&gt;AWS has &lt;a href="https://aws.amazon.com/about-aws/whats-new/2023/11/amazon-sqs-throughput-quota-fifo-high-throughput-mode/" rel="noopener noreferrer"&gt;dramatically increased the throughput for SQS FIFO&lt;/a&gt;, now enabling processing of up to 70,000 messages per second in &lt;a href="https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/high-throughput-fifo.html#enable-high-throughput-fifo" rel="noopener noreferrer"&gt;high throughput mode&lt;/a&gt;. This enhancement marks a significant leap in handling large volumes of messages efficiently, catering to more demanding and high-traffic applications.&lt;/p&gt;

&lt;h2&gt;
  
  
  Aurora Limitless Database
&lt;/h2&gt;

&lt;p&gt;Amazon Aurora has launched the &lt;a href="https://aws.amazon.com/about-aws/whats-new/2023/11/amazon-aurora-limitless-database/" rel="noopener noreferrer"&gt;Aurora Limitless Database&lt;/a&gt;, a significant upgrade allowing clusters to scale up to millions of write transactions per second and manage petabytes of data. While most users may not require this extreme level of scalability, the technical achievement is impressive and noteworthy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final thoughts
&lt;/h2&gt;

&lt;p&gt;re:Invent 2023 showcased an array of remarkable advancements, particularly in the Serverless domain. From the significant scaling improvements in Lambda and SQS FIFO to the innovative features in Step Functions and the technical prowess of Aurora Limitless, AWS is pushing the boundaries of what's possible in cloud computing. While some offerings, like ElastiCache Serverless, sparked debate over pricing and operational aspects, the overall direction is clear: AWS is committed to providing more robust, scalable, and efficient solutions, driving the future of Serverless computing forward. As we embrace these changes, it's exciting to ponder how they will shape our technological landscape in the coming years.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>serverless</category>
    </item>
  </channel>
</rss>
