Most MCP servers get tested the same way: connect one client, call a few tools by hand, ship it.
Then real agents show up. Dozens of sessions at once, three tool calls in parallel, sessions that never get closed. The failures that follow don't look like crashes. They look like memory that creeps up for hours, a p95 that slowly doubles, or Session not found errors that only appear once you run two replicas behind a load balancer.
This post shows how to find those problems on your own machine in under an hour, using mcpload, an open-source (Apache-2.0) load and soak tester for MCP servers built on k6.
As the example I'll use the official MCP reference server, server-everything, and share what I measured.
What a "soak test" actually checks
A load test asks how much can it take? A soak test asks does it stay healthy over time?
mcpload's soak run has three phases:
- Warm-up: load ramps up. Startup growth (caches, connection pools) is allowed here.
- Steady load: a constant stream of agent sessions for 30 minutes.
- Cool-down: no load for 5 minutes. Memory should come back down.
A leak is flagged only if memory grows steadily during the load window (slope above a limit, with a good linear fit) or doesn't come back near its baseline during cool-down. "Memory went up" alone is not a leak. Most servers grow a bit at startup, and treating that as a leak causes false alarms.
Each simulated agent does what a real one does: initialize → tools/list → 1–5 rounds of 3 parallel tools/call with think time → close the session.
Step 1: Install mcpload
Download the release for your platform from GitHub Releases. Linux example:
curl -fL https://github.com/atul121001/mcpload/releases/download/v0.1.2/mcpload_0.1.2_linux_amd64.tar.gz | tar -xz
cd mcpload_0.1.2_linux_amd64
./mcpload version
Mac: use darwin_arm64 (Apple silicon) or darwin_amd64 (Intel). Windows: download the .zip.
The folder contains mcpload (the CLI), a k6 binary with MCP support built in, and the bundled test scenarios.
Step 2: Start the server in Docker
Running the server in a container gives it fixed resources, and lets mcpload read its memory from Docker:
docker run -d --name mcp-under-test -p 127.0.0.1:5101:5101 -e PORT=5101 \
--cpus 2 --memory 1g \
node:22-slim npx -y @modelcontextprotocol/server-everything@2026.8.31 streamableHttp
Wait until docker logs mcp-under-test shows listening on port 5101.
Step 3: A one-minute smoke test
./mcpload run --url http://localhost:5101/mcp --duration 1m \
--env 'TOOL_MIX={"echo":3,"get-sum":2,"get-tiny-image":1}' \
--env 'TOOL_ARGS={"echo":{"message":"hello"},"get-sum":{"a":2,"b":3}}'
TOOL_MIX picks which tools to call and how often. Only include tools that are safe to call thousands of times: no tools that send email, write to production, or call paid APIs. Without it, mcpload calls every tool the server lists, with arguments generated from each tool's schema.
You should see mcpload result: PASS.
Step 4: The 38-minute soak
./mcpload run --url http://localhost:5101/mcp --scenario soak \
--soak-min 30 --warmup-min 3 --cooldown-min 5 --env RATE=2 \
--env 'TOOL_MIX={"echo":3,"get-sum":2,"get-tiny-image":1}' \
--env 'TOOL_ARGS={"echo":{"message":"hello"},"get-sum":{"a":2,"b":3}}' \
--sampler docker --container mcp-under-test \
--out soak.json --html soak.html
RATE=2 means two new agent sessions start every second, about 3,600 sessions over the 30 minutes. --sampler docker records the container's memory every 10 seconds.
What I measured
On my laptop (container limited to 2 CPUs and 1 GiB), server-everything 2026.8.31 (TypeScript SDK 1.31):
- 3,780 sessions, 49,227 requests, 0 errors
- Memory flat: ~136 MiB, growing 0.12 MiB/min against a 1 MiB/min limit, and back near baseline after cool-down
- p95 latency steady at ~21 ms for all three tools over the full 30 minutes
[screenshot: soak.html memory chart]
I also stepped the load from 5 to 80 concurrent agents, 2 minutes each:
| Concurrent agents | Tool calls/s | p95 |
|---|---|---|
| 5 | 40 | 23 ms |
| 10 | 81 | 21 ms |
| 20 | 172 | 22 ms |
| 40 | 333 | 25 ms |
| 80 | 534 | 178 ms |
Throughput scaled linearly up to 40 agents. At 80, latency rose sharply but there were still zero errors. On 2 cores, the knee for this workload lies between 40 and 80 concurrent agents.
That's a healthy server. The point of the exercise is that you now know where its limits are before your users find them.
These are numbers from one laptop run of a demo server with trivial tools. They're not a benchmark, and they don't say anything about production deployments.
Things worth testing on your own server
-
Behind a load balancer. If your server uses stateful sessions (
Mcp-Session-Id), run two replicas without sticky sessions and use--scenario lb-check. mcpload countssession_not_founderrors separately. - Sessions that never close. Agents crash and drop connections. If your server keeps per-session state, the soak's cool-down shows whether it's ever freed.
-
Bursts.
--scenario burstfloodsinitializeand then ramps to 200 agents at once.
Run it in CI
There's a GitHub Action that runs the test on every pull request and posts a per-tool table as a PR comment:
- uses: atul121001/mcpload-action@v1
with:
url: http://localhost:8080/mcp
vus: '10'
duration: 2m
p95-ms: '800'
What's next
I'm running the same test across the official MCP SDKs (Python, Go, Rust, C#), and I'll share the results with each maintainer before publishing.
If you try mcpload on your own server, I'd love to hear what it finds, or what it should check that it doesn't yet. Repo: https://github.com/atul121001/mcpload
Top comments (2)
The tricky failure mode that clean echo tools hide is child process cleanup on agent cancellation. When an agent times out or crashes midway through a long-running tool call, the HTTP transport tears down the session cleanly, but the server often leaves the underlying ripgrep or shell worker running in the background. After an hour under twenty concurrent agents, memory looks stable on paper while the host runs out of file descriptors and zombie process slots. Adding an explicit process tree kill on transport disconnect saved our runners.
This is a great one, and you're right that a clean echo server hides it completely. mcpload doesn't catch it today: it only abandons a call on timeout, and the Docker sampler tracks memory, not processes or file descriptors.
Two things I'm adding because of your comment: a mode that cancels a share of tool calls mid-flight (client disconnects before the result), and process + fd counts from the container so a pid_leak / fd_leak verdict can fire when workers outlive their sessions. Process-tree kill on disconnect is a great fix to point people to as well.
Out of curiosity, was it mostly ripgrep / shell tools, or long-running subprocesses in general?