Most MCP demos show a tool that answers in a second. Real tools often don't: a scraper, a crawl or a report job can take minutes. We build web scrapers on Apify, so this is a problem we run into daily. We wanted to know what an agent actually receives when a remote MCP tool runs longer than one call, so we traced the raw JSON-RPC traffic against Apify's hosted MCP server (mcp.apify.com, version 0.17.1 at the time) on Oct 2, 2026. The job: a YellowPages scraper reading one results page, which takes about 35 seconds.
1. Discovery: one tool per job, plus helpers
After initialize, tools/list returned one tool per scraper we had allowed through the server URL, plus four helper tools that only make sense for long-running work:
-
get-actor-runchecks a run's status and can wait for it, -
get-dataset-itemsreads the results in pages, -
get-key-value-store-recordreads a stored file or record, -
abort-actor-runstops a run.
The helpers are the first hint: the server expects the agent to come back for results.
2. The call returns early
The agent calls the scraper tool with ordinary arguments:
{ "searchKeyword": ["plumbers"], "location": "Austin, TX", "maxPages": 1 }
The response arrived after about 30 seconds, while the job was still running. It had two content parts. The first is structured:
{
"runId": "…",
"status": "RUNNING",
"storages": { "datasets": { "default": { "id": "…", "itemCount": 23 } } }
}
The second is plain text written for the model:
RUNNING for 30s. In progress. 23 results so far.
Use get-actor-run with runId=… and waitSecs=30 to poll for completion.
That second part is the interesting design choice. The server returns machine-readable state and also tells the model, in plain English, what to do next. An agent that only parses the JSON still works. An agent that only reads text also knows to poll.
3. Polling has a ceiling
We first asked get-actor-run to wait 60 seconds and then 120. Both came back as a JSON-RPC error, not a tool result:
MCP error -32602: Invalid arguments for tool "get-actor-run".
Validation errors: /waitSecs: must be <= 45.
With waitSecs: 30 the call returned SUCCEEDED, the run time (about 35 seconds) and the dataset item count (37). The result also carried a _meta block with the run's platform usage, which a client can show the user before it starts anything bigger.
Two things to handle here:
- Validation problems arrive as protocol errors (
-32602), not as tool results withisError. If your client only checksisError, it will miss them. - The wait is capped, so a five-minute job means several polls. Your loop needs a total deadline, not just a per-call wait.
4. Reading results in pages
get-dataset-items takes the dataset ID plus offset and limit, and returns:
{ "datasetId": "…", "items": [ … ], "itemCount": 37, "totalItemCount": 37, "offset": 0, "limit": 100 }
For a big job, page through with offset instead of pulling everything into the model's context at once.
5. The trap: SUCCEEDED with zero items
The same evening a second test job, a Google Maps scraper, ran for about 157 seconds and finished with status SUCCEEDED, but get-dataset-items came back empty:
{ "items": [], "itemCount": 0, "totalItemCount": 0, "offset": 0, "limit": 100 }
An upstream dependency was failing, and the job ended cleanly with nothing to show. If your agent treats SUCCEEDED as "done, report success", it will tell the user everything worked. Check the item count, and treat "succeeded but empty" as a result the user needs to hear about.
A checklist for long-running MCP tools
- Expect an early return. Store the run ID the first response gives you.
- Poll within the server's limits (45 seconds per wait here), with an overall deadline and a clear message if it runs out.
-
Check protocol errors and
isError. They come back in different places. - Read results in pages to protect the context window.
- Treat "succeeded, zero items" as a warning, not a success.
-
Abort what you no longer need. If the user changes the question halfway,
abort-actor-runsaves the rest of the job's cost. - Expose fewer tools. A server-side allow list keeps the model from picking the wrong tool and limits what it can touch.
How does your MCP client deal with tools that outlive one call: polling, progress notifications, or something else? We're curious what holds up in production.
Written with AI assistance. Every request and response above comes from real calls we made on Oct 2, 2026.
Top comments (0)