DEV Community

jidonglab
jidonglab

Posted on

Claude Parallel Tool Use: My Agent Loop Crashed on 388 of 1,200 Runs

My codebase Q&A agent ran 1,200 times in its first week. 388 of those runs died with the same HTTP 400. The other 812 worked fine, which is the worst possible ratio: too low to look like a broken deploy, too high to ignore.

The cause was one line I'd copied from my own earlier prototype: next(b for b in resp.content if b.type == "tool_use"). It grabs the first tool call. Claude parallel tool use means there often isn't just one.

Then I "fixed" it in a way that made the crashes disappear and quietly made the agent dumber. This post covers all of it, plus the numbers from the version that finally worked.

TL;DR

  • With Claude parallel tool use, one assistant turn can contain several tool_use blocks. In my logs, 14.2% of tool turns had 2 or more, up to 6.
  • Every tool_use id must get a matching tool_result in the very next user message, or the API returns a 400.
  • Do not strip the extra tool_use blocks to silence the error. It stops the crash and degrades answers (my eval went from 46/50 to 38/50).
  • The fix: execute every call (concurrently with asyncio.gather), return all results in one user message, and send failures as is_error: true instead of dropping them.
  • Only set disable_parallel_tool_use when tool order truly matters. It cost me 0.8 extra turns per run.

What does the agent actually do?

It answers questions about a mid-sized Python monorepo. Three tools: read_file, grep, list_dir. Python SDK, a Claude Sonnet model, a plain while stop_reason == "tool_use" loop. Nothing exotic.

A typical question is "where do we retry failed webhook deliveries?" The model greps, reads two or three files, answers. On a good run that's 3 or 4 round trips.

Here's the loop body I shipped on day one:

resp = client.messages.create(model=MODEL, max_tokens=2048,
                              tools=TOOLS, messages=messages)

if resp.stop_reason == "tool_use":
    block = next(b for b in resp.content if b.type == "tool_use")
    result = run_tool(block.name, block.input)
    messages.append({"role": "assistant", "content": resp.content})
    messages.append({"role": "user", "content": [
        {"type": "tool_result", "tool_use_id": block.id, "content": result}
    ]})
Enter fullscreen mode Exit fullscreen mode

It passed every manual test I ran. My manual tests were all simple questions, and simple questions get one tool call at a time.

Why does Claude return multiple tool_use blocks in one response?

Because it can, and it's usually the smart move. When the model already knows it needs retry.py and webhooks/sender.py, asking for both in one turn saves a full round trip. Parallel tool use is on by default in the Messages API.

So a response to "compare how the two webhook senders handle timeouts" looks like this:

content: [
  text:     "I'll read both sender implementations."
  tool_use: id=toolu_01A..., name=read_file, input={path: "webhooks/sender.py"}
  tool_use: id=toolu_01B..., name=read_file, input={path: "webhooks/legacy_sender.py"}
]
stop_reason: "tool_use"
Enter fullscreen mode Exit fullscreen mode

My loop ran toolu_01A, appended the full assistant content (both blocks), and sent back one result. The next request failed with:

400 invalid_request_error: messages.4: `tool_use` ids were found without
`tool_result` blocks immediately after: toolu_01B...
Enter fullscreen mode Exit fullscreen mode

That's the rule: every tool_use in an assistant message needs a tool_result with the matching tool_use_id in the next user message. No partial credit.

How did the "fix" make my agent worse?

My first patch was the obvious shortcut. If the API complains about orphaned tool_use blocks, remove them before appending the assistant turn:

first = next(b for b in resp.content if b.type == "tool_use")
kept = [b for b in resp.content if b.type != "tool_use" or b.id == first.id]
messages.append({"role": "assistant", "content": kept})
Enter fullscreen mode Exit fullscreen mode

Crashes went to zero. I felt great for about a day.

Then I ran my 50-question eval set (hand-labeled, each with a known correct file and line range). Score: 38/50. The v4 loop described below scores 46/50 on the same set.

Reading transcripts made it obvious. From the model's point of view, its own history now said it had asked for one file. So it did one of two things:

  1. Asked again. It re-requested the files it "forgot," burning extra turns. Average turns per run rose to 4.9.
  2. Answered without them. Sometimes it decided one file was enough and confidently described the legacy sender using code it had never seen. Those were the wrong answers.

Editing the assistant's own past turns is gaslighting your agent. It plans based on what it believes it already did.

What is the correct way to handle Claude parallel tool calls?

Execute every tool_use block, then return every result together in a single user message. Here's the loop I run now:

import asyncio
from anthropic import AsyncAnthropic

client = AsyncAnthropic()

async def run_one(block):
    try:
        out = await run_tool(block.name, block.input)
        return {"type": "tool_result", "tool_use_id": block.id, "content": out}
    except Exception as e:
        return {"type": "tool_result", "tool_use_id": block.id,
                "content": f"{type(e).__name__}: {e}", "is_error": True}

async def step(messages):
    resp = await client.messages.create(model=MODEL, max_tokens=2048,
                                        tools=TOOLS, messages=messages)
    messages.append({"role": "assistant", "content": resp.content})
    if resp.stop_reason != "tool_use":
        return resp

    calls = [b for b in resp.content if b.type == "tool_use"]
    results = await asyncio.gather(*(run_one(b) for b in calls))

    assert {r["tool_use_id"] for r in results} == {c.id for c in calls}
    messages.append({"role": "user", "content": list(results)})
    return None
Enter fullscreen mode Exit fullscreen mode

Three details matter more than they look.

One user message, not several. I briefly tried sending one user message per result. Beyond being awkward, Anthropic's tool use docs warn that splitting results across messages teaches the model to stop making parallel calls in that conversation. You lose the speedup you just paid to support.

tool_result blocks go first. If you want to add text to that user message (I append a short "N tool calls remaining in budget" note), it goes after the results. Text first and the API rejects it.

Errors are results. A FileNotFoundError on one of three reads used to kill the whole turn. Now the model gets is_error: true with the message, and it usually corrects the path on the next turn. That alone recovered 2 of my 50 eval questions.

The assert is cheap insurance. If someone later adds a tool that silently returns None and gets filtered out, the loop fails loudly in my process instead of with a 400 from the API.

How often does Claude actually make parallel tool calls?

In my setup, about one tool turn in seven. After the fix, I logged every turn across another 1,200 runs:

  • 4,310 assistant turns ended with stop_reason: "tool_use"
  • 611 of them (14.2%) contained two or more tool_use blocks
  • the largest single turn had 6 (read_file across six test fixtures)
  • grep + read_file pairs were rare; parallel turns were almost always same-tool fan-outs

That 14.2% is the number that explains the 32% crash rate. A run has several tool turns, so the odds that at least one of them goes parallel are much higher than the per-turn rate. Your number will differ with your tools and prompts. Read-heavy tools with obvious fan-out invite more parallelism than a single "run SQL" tool would.

Should I just set disable_parallel_tool_use?

Only if your tools have ordering dependencies. Setting tool_choice={"type": "auto", "disable_parallel_tool_use": True} makes the model emit at most one tool_use per turn, which makes the naive loop technically correct.

I tested it as a fourth variant on the same question set. All four, side by side:

Loop version Crashed runs Eval score Turns/run Input tokens/run Median latency
v1: first block only 388 / 1,200 n/a n/a n/a n/a
v2: strip extra blocks 0 38/50 4.9 ~41K 26.5s
v3: disable_parallel_tool_use 0 45/50 4.4 ~36K 24.9s
v4: all results, gathered 0 46/50 3.6 ~29K 18.2s

The token column surprised me most. Every extra turn re-sends the whole conversation as input. Fewer round trips means less re-reading, so v4 used about 29% fewer input tokens per run than v2, on top of being faster because the file reads overlap.

Where I would turn parallelism off: an agent with write_file and run_tests. If the model fires both in one turn, asyncio.gather gives you no ordering guarantee, and tests may run against the old file. For side-effecting tools, either disable parallel use or execute the blocks sequentially in the order they appear.

What should I check in my own agent loop?

Grep your code for these patterns. Each one is a version of my bug:

  • next(b for b in ... if b.type == "tool_use")
  • resp.content[-1] or resp.content[1] used as "the tool call"
  • a tool_use_id variable that's a single string instead of a list
  • any try/except around tool execution that continues without appending a result
  • code that edits or filters resp.content before appending it to history

If your SDK version ships a tool runner helper, it handles the matching for you. I still like owning the loop, because then I can log the parallel rate, which turned out to be the most useful number in this whole debugging session.

The short answer

Claude parallel tool use returns multiple tool_use blocks in a single response, and the API requires a tool_result for every one of them, matched by tool_use_id, in the next user message. Loops that handle only the first block crash with a 400; loops that strip the extra blocks stop crashing but feed the model a false history and give worse answers. Execute every call, return all results together in one user message with failures marked is_error: true, and reserve disable_parallel_tool_use for tools whose order matters. In my agent that meant zero crashes, 46/50 on eval instead of 38/50, and runs that were 31% faster.


Written by the developer behind Preterview, an interview prep platform.

Top comments (1)

Collapse
 
mateo_ruiz_6992b1fce47843 profile image
Mateo Ruiz •

The most important lesson here is that tool-call handling is part of the agent’s state machine, not just API plumbing. Once the model can emit multiple tool_use blocks, the loop has to preserve the exact relationship between intent, execution, and result.

This is something we pay close attention to at IT Path Solutions when building production agent workflows: a failed tool call shouldn't erase the model's observation that the call was attempted. Returning is_error: true preserves that history and gives the agent a chance to recover instead of forcing the whole turn to restart.

I’d also make one distinction explicit: parallelism is safe when the tools are observational and independent; it becomes a correctness problem as soon as tools mutate shared state. That makes “can these calls run concurrently?” a property of the tool contract, not simply a model configuration.

The 14.2% parallel-turn rate is a useful metric for exactly that reason. Logging it alongside tool latency, failure rate, and retry behavior gives you a much better view of whether parallel execution is actually improving the workflow or just exposing hidden ordering assumptions.