<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Atul Mishra</title>
    <description>The latest articles on DEV Community by Atul Mishra (@atulmishra).</description>
    <link>https://dev.to/atulmishra</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F767329%2F98a17f97-bb9c-41e4-a8b2-357d174d8d07.png</url>
      <title>DEV Community: Atul Mishra</title>
      <link>https://dev.to/atulmishra</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/atulmishra"/>
    <language>en</language>
    <item>
      <title>How to soak-test your MCP server before AI agents do it for you</title>
      <dc:creator>Atul Mishra</dc:creator>
      <pubDate>Fri, 02 Oct 2026 18:55:18 +0000</pubDate>
      <link>https://dev.to/atulmishra/how-to-soak-test-your-mcp-server-before-ai-agents-do-it-for-you-2jc3</link>
      <guid>https://dev.to/atulmishra/how-to-soak-test-your-mcp-server-before-ai-agents-do-it-for-you-2jc3</guid>
      <description>&lt;p&gt;Most MCP servers get tested the same way: connect one client, call a few tools by hand, ship it.&lt;/p&gt;

&lt;p&gt;Then real agents show up. Dozens of sessions at once, three tool calls in parallel, sessions that never get closed. The failures that follow don't look like crashes. They look like memory that creeps up for hours, a p95 that slowly doubles, or &lt;code&gt;Session not found&lt;/code&gt; errors that only appear once you run two replicas behind a load balancer.&lt;/p&gt;

&lt;p&gt;This post shows how to find those problems on your own machine in under an hour, using &lt;strong&gt;&lt;a href="https://github.com/atul121001/mcpload" rel="noopener noreferrer"&gt;mcpload&lt;/a&gt;&lt;/strong&gt;, an open-source (Apache-2.0) load and soak tester for MCP servers built on &lt;a href="https://k6.io" rel="noopener noreferrer"&gt;k6&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;As the example I'll use the official MCP reference server, &lt;code&gt;server-everything&lt;/code&gt;, and share what I measured.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a "soak test" actually checks
&lt;/h2&gt;

&lt;p&gt;A load test asks &lt;em&gt;how much can it take?&lt;/em&gt; A soak test asks &lt;em&gt;does it stay healthy over time?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;mcpload's soak run has three phases:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Warm-up&lt;/strong&gt;: load ramps up. Startup growth (caches, connection pools) is allowed here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Steady load&lt;/strong&gt;: a constant stream of agent sessions for 30 minutes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cool-down&lt;/strong&gt;: no load for 5 minutes. Memory should come back down.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A leak is flagged only if memory grows steadily during the load window (slope above a limit, with a good linear fit) &lt;strong&gt;or&lt;/strong&gt; doesn't come back near its baseline during cool-down. "Memory went up" alone is not a leak. Most servers grow a bit at startup, and treating that as a leak causes false alarms.&lt;/p&gt;

&lt;p&gt;Each simulated agent does what a real one does: &lt;code&gt;initialize&lt;/code&gt; → &lt;code&gt;tools/list&lt;/code&gt; → 1–5 rounds of 3 parallel &lt;code&gt;tools/call&lt;/code&gt; with think time → close the session.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: Install mcpload
&lt;/h2&gt;

&lt;p&gt;Download the release for your platform from &lt;a href="https://github.com/atul121001/mcpload/releases/latest" rel="noopener noreferrer"&gt;GitHub Releases&lt;/a&gt;. Linux example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fL&lt;/span&gt; https://github.com/atul121001/mcpload/releases/download/v0.1.2/mcpload_0.1.2_linux_amd64.tar.gz | &lt;span class="nb"&gt;tar&lt;/span&gt; &lt;span class="nt"&gt;-xz&lt;/span&gt;
&lt;span class="nb"&gt;cd &lt;/span&gt;mcpload_0.1.2_linux_amd64
./mcpload version
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Mac: use &lt;code&gt;darwin_arm64&lt;/code&gt; (Apple silicon) or &lt;code&gt;darwin_amd64&lt;/code&gt; (Intel). Windows: download the &lt;code&gt;.zip&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The folder contains &lt;code&gt;mcpload&lt;/code&gt; (the CLI), a &lt;code&gt;k6&lt;/code&gt; binary with MCP support built in, and the bundled test scenarios.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: Start the server in Docker
&lt;/h2&gt;

&lt;p&gt;Running the server in a container gives it fixed resources, and lets mcpload read its memory from Docker:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; mcp-under-test &lt;span class="nt"&gt;-p&lt;/span&gt; 127.0.0.1:5101:5101 &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="nv"&gt;PORT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;5101 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cpus&lt;/span&gt; 2 &lt;span class="nt"&gt;--memory&lt;/span&gt; 1g &lt;span class="se"&gt;\&lt;/span&gt;
  node:22-slim npx &lt;span class="nt"&gt;-y&lt;/span&gt; @modelcontextprotocol/server-everything@2026.8.31 streamableHttp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Wait until &lt;code&gt;docker logs mcp-under-test&lt;/code&gt; shows &lt;code&gt;listening on port 5101&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: A one-minute smoke test
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;./mcpload run &lt;span class="nt"&gt;--url&lt;/span&gt; http://localhost:5101/mcp &lt;span class="nt"&gt;--duration&lt;/span&gt; 1m &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--env&lt;/span&gt; &lt;span class="s1"&gt;'TOOL_MIX={"echo":3,"get-sum":2,"get-tiny-image":1}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--env&lt;/span&gt; &lt;span class="s1"&gt;'TOOL_ARGS={"echo":{"message":"hello"},"get-sum":{"a":2,"b":3}}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;TOOL_MIX&lt;/code&gt; picks which tools to call and how often. Only include tools that are safe to call thousands of times: no tools that send email, write to production, or call paid APIs. Without it, mcpload calls every tool the server lists, with arguments generated from each tool's schema.&lt;/p&gt;

&lt;p&gt;You should see &lt;code&gt;mcpload result: PASS&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: The 38-minute soak
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;./mcpload run &lt;span class="nt"&gt;--url&lt;/span&gt; http://localhost:5101/mcp &lt;span class="nt"&gt;--scenario&lt;/span&gt; soak &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--soak-min&lt;/span&gt; 30 &lt;span class="nt"&gt;--warmup-min&lt;/span&gt; 3 &lt;span class="nt"&gt;--cooldown-min&lt;/span&gt; 5 &lt;span class="nt"&gt;--env&lt;/span&gt; &lt;span class="nv"&gt;RATE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;2 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--env&lt;/span&gt; &lt;span class="s1"&gt;'TOOL_MIX={"echo":3,"get-sum":2,"get-tiny-image":1}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--env&lt;/span&gt; &lt;span class="s1"&gt;'TOOL_ARGS={"echo":{"message":"hello"},"get-sum":{"a":2,"b":3}}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--sampler&lt;/span&gt; docker &lt;span class="nt"&gt;--container&lt;/span&gt; mcp-under-test &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--out&lt;/span&gt; soak.json &lt;span class="nt"&gt;--html&lt;/span&gt; soak.html
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;RATE=2&lt;/code&gt; means two new agent sessions start every second, about 3,600 sessions over the 30 minutes. &lt;code&gt;--sampler docker&lt;/code&gt; records the container's memory every 10 seconds.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I measured
&lt;/h2&gt;

&lt;p&gt;On my laptop (container limited to 2 CPUs and 1 GiB), &lt;code&gt;server-everything&lt;/code&gt; 2026.8.31 (TypeScript SDK 1.31):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;3,780 sessions, 49,227 requests, 0 errors&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory flat&lt;/strong&gt;: ~136 MiB, growing 0.12 MiB/min against a 1 MiB/min limit, and back near baseline after cool-down&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;p95 latency steady at ~21 ms&lt;/strong&gt; for all three tools over the full 30 minutes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;[screenshot: soak.html memory chart]&lt;/p&gt;

&lt;p&gt;I also stepped the load from 5 to 80 concurrent agents, 2 minutes each:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Concurrent agents&lt;/th&gt;
&lt;th&gt;Tool calls/s&lt;/th&gt;
&lt;th&gt;p95&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;23 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;81&lt;/td&gt;
&lt;td&gt;21 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;172&lt;/td&gt;
&lt;td&gt;22 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;40&lt;/td&gt;
&lt;td&gt;333&lt;/td&gt;
&lt;td&gt;25 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;80&lt;/td&gt;
&lt;td&gt;534&lt;/td&gt;
&lt;td&gt;178 ms&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Throughput scaled linearly up to 40 agents. At 80, latency rose sharply but there were still zero errors. On 2 cores, the knee for this workload lies between 40 and 80 concurrent agents.&lt;/p&gt;

&lt;p&gt;That's a healthy server. The point of the exercise is that you now know &lt;em&gt;where&lt;/em&gt; its limits are before your users find them.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;These are numbers from one laptop run of a demo server with trivial tools. They're not a benchmark, and they don't say anything about production deployments.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Things worth testing on your own server
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Behind a load balancer.&lt;/strong&gt; If your server uses stateful sessions (&lt;code&gt;Mcp-Session-Id&lt;/code&gt;), run two replicas without sticky sessions and use &lt;code&gt;--scenario lb-check&lt;/code&gt;. mcpload counts &lt;code&gt;session_not_found&lt;/code&gt; errors separately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sessions that never close.&lt;/strong&gt; Agents crash and drop connections. If your server keeps per-session state, the soak's cool-down shows whether it's ever freed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bursts.&lt;/strong&gt; &lt;code&gt;--scenario burst&lt;/code&gt; floods &lt;code&gt;initialize&lt;/code&gt; and then ramps to 200 agents at once.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Run it in CI
&lt;/h2&gt;

&lt;p&gt;There's a GitHub Action that runs the test on every pull request and posts a per-tool table as a PR comment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;atul121001/mcpload-action@v1&lt;/span&gt;
  &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;url&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://localhost:8080/mcp&lt;/span&gt;
    &lt;span class="na"&gt;vus&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;10'&lt;/span&gt;
    &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;2m&lt;/span&gt;
    &lt;span class="na"&gt;p95-ms&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;'&lt;/span&gt;&lt;span class="s"&gt;800'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;I'm running the same test across the official MCP SDKs (Python, Go, Rust, C#), and I'll share the results with each maintainer before publishing.&lt;/p&gt;

&lt;p&gt;If you try mcpload on your own server, I'd love to hear what it finds, or what it should check that it doesn't yet. Repo: &lt;strong&gt;&lt;a href="https://github.com/atul121001/mcpload" rel="noopener noreferrer"&gt;https://github.com/atul121001/mcpload&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>testing</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
