DEV Community

Cover image for Claude Code's Background-Agent Wait Ceiling: A $4.19 Lesson in Headless Design
soy
soy

Posted on Originally published at media.patentllm.org

Claude Code's Background-Agent Wait Ceiling: A $4.19 Lesson in Headless Design

A headless Claude Code pipeline that fans work out to background subagents can appear to succeed, exit code 0, a clean JSON result, while producing zero output and burning real API cost. This repository's own cron-driven news pipeline hit that failure mode once, and the fix was a single environment variable most headless setups never touch.

How it works

scripts/agent_news.sh drives the daily news pipeline as a single headless call: claude -p "$PROMPT" --model claude-sonnet-5 --allowedTools Read Write Edit Bash WebFetch WebSearch Glob Grep --permission-mode acceptEdits --output-format json. The prompt tells the agent to follow scripts/agent_news.md, which instructs it to fetch roughly 51 RSS/Atom/changelog sources in parallel, since fetching them one at a time takes too long. In practice, that instruction gets carried out by spawning multiple background Task subagents, one per group of sources, rather than issuing dozens of foreground tool calls inside a single turn.

On 2026-09-15 18:00 JST, that run spawned 7 background subagents. The harness has a default wait ceiling on how long a claude -p invocation will block on outstanding background work; six of the seven subagents finished, the seventh was still running when the ceiling was reached, and the harness terminated it (subagent_stats.killed.system: 1). The parent process then returned normally, exit_code=0, a well-formed JSON result, with the model's own final message reading only '最後のグループ(hardware B + Google系)の完了を待っています' ('still waiting on the last group'). Zero articles were saved. The run had already spent $4.19 (haiku-4-5 subagent calls: $1.67; sonnet-5 orchestration: $2.52) by the time it gave up.

The fix, applied the same day, was to stop relying on the default wait ceiling and pin it explicitly:

# 0 にすると無制限に待つので使わない。上限そのものが停止の歯止めになる。
export CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS="${CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS:-1800000}"
Enter fullscreen mode Exit fullscreen mode

30 minutes (1,800,000 ms) instead of the harness's implicit 600,000 ms default. A manual re-run at 18:22:45 that day spawned 5 background subagents, all of which completed inside the new window, and produced all 5 category articles plus the Deep Dive, at a cost of $5.76 for the successful run, on top of the $4.19 already spent on the failed one.

When this helps

Anyone running claude -p headless with tool access to spawn background Task subagents, a common pattern for fanning out independent fetches, searches, or file scans, should pin CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS explicitly rather than trusting the harness default. The risk is specific to background subagents: a single foreground turn that calls many tools in parallel (as this same task did today, fetching over 50 sources directly inside one turn instead of spawning subagents) does not hit this ceiling at all, because there is no separate background wait to time out. The ceiling only matters once background Task calls enter the picture, nightly cron jobs, long-running research fan-outs, or any headless pipeline where a human isn't watching the terminal to notice a stall. Interactive sessions are lower risk: a stuck background agent is visible in the UI, and a human can wait, kill it, or retry immediately, instead of an exit=0 silently reaching a notification script that only checks the log for known failure markers, not whether the ceiling was hit.

Measured here

  • 2026-09-15 18:00 JST(タイムアウト失敗): background subagent 7起動・6完了・1件がシステムによりkill、exit=0、記事0件保存、コスト$4.19(claude-haiku-4-5: $1.67 / claude-sonnet-5: $2.52)
  • 同日18:22:45(CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=1800000設定後の再実行): background subagent 5起動・5完了、記事5カテゴリ+Deep Dive保存、コスト$5.76(claude-haiku-4-5: $1.83 / claude-sonnet-5: $3.93)
  • ハーネスの既定の待機上限は600,000ms(10分)。ログの実測メッセージ「Background tasks still running after 600s; terminating.」より。

Is it worth it

Setting the wait ceiling costs nothing and pays for itself the first time a background fan-out runs long: raising it from the harness default to 30 minutes turned a guaranteed-timeout failure into a normal, billable success on this pipeline's very next attempt. The harder lesson is architectural, not the env var itself, deciding whether a multi-source fetch belongs in background subagents at all. Today's run of this same procedure used direct, foreground parallel tool calls instead of Task subagents for the source fetch, and never touched the ceiling. Background subagents earn their keep when the work needs an isolated context window or must keep running after the orchestrating turn ends; for a bounded, known list of sources fetched once, foreground parallel calls remove the timeout risk entirely rather than just widening the window it can fail in. Worth adopting either way: pin the ceiling if background subagents are in use, and ask first whether they need to be.

Written while running Claude Code unattended in production every day. Full archive: https://media.patentllm.org

Top comments (0)