DEV Community

Cover image for Claude Headless Sandboxing: Why /tmp Writes Fail in Non-Interactive Mode
Guatu
Guatu

Posted on Originally published at guatulabs.dev

Claude Headless Sandboxing: Why /tmp Writes Fail in Non-Interactive Mode

A scheduled Claude Code job that writes its report to /tmp/agent_log.md works every time you run it by hand. Under a systemd timer with claude -p, it exits 0, the journal shows no errors, and the file is either missing or empty.

That's the whole symptom. There's no stack trace to search for, which is why this one costs people an afternoon.

I hit this with my second-brain curator. It's a shell script on a timer that points Claude at a pile of session notes, has it pick out what's worth promoting into long-term memory, and writes a curation report. That's the pipeline behind the self-updating wiki setup. The fix itself was small: move the logs out of /tmp and into $HOME/logs. The reason the fix works is the interesting part. It says a lot about how headless agents see the filesystem, and about how interactive mode hides the permission model from you.

What I expected

My mental model was plain Unix. The script runs as my user, /tmp is world-writable with the sticky bit, so any process running as me can create files there. Claude Code is a process running as me. Therefore Claude Code can write to /tmp.

Interactive sessions supported that model. I'd ask Claude to "write the summary to /tmp/curator-test.md" and it would. The file showed up and I moved on.

The broken version looked something like this:

#!/usr/bin/env bash
# curator.sh (the version that silently failed under systemd)
set -euo pipefail

claude -p "Review the notes in ./inbox. Write a curation report \
to /tmp/agent_log.md listing which notes to promote and why." \
  --allowedTools "Read,Glob,Grep,Write"

echo "curator finished"
Enter fullscreen mode Exit fullscreen mode

Run from a terminal, this worked, and the reason is the whole problem.

What actually happened

Two separate mechanisms were stacked on top of each other, and neither of them has anything to do with Unix permissions on /tmp.

1. Claude Code's permission boundary is the working directory, not your UID

Claude Code doesn't scope file access by what your user account can touch. It scopes it by working directories: the directory you launched from, plus anything you explicitly add. File edits inside that boundary follow your permission mode. File edits outside it need approval.

In an interactive session you rarely notice, because when Claude tries to write /tmp/curator-test.md you get a permission prompt, hit "yes," and forget it happened. Maybe you picked "don't ask again" and it went into your local settings, and now you've forgotten that too.

In print mode (-p) nobody is there to answer the prompt. The tool call gets denied, and the model is told the write wasn't allowed. Then the model does what models do: it adapts. Sometimes it prints the report to stdout instead. Sometimes it writes a note saying it couldn't save the file. Sometimes it just summarizes what it would have written. The process finishes normally and exits 0.

So from the script's point of view, nothing failed. An agent that works around a denied permission looks like success to a shell script. That's the gap.

Putting Write in --allowedTools doesn't get you past the boundary either. Allowing the tool isn't the same as allowing the tool to reach any path. The directory scoping still applies.

2. systemd has its own opinions about /tmp

Even after fixing the permission boundary, /tmp under systemd is its own trap. If the service unit (or a drop-in you've forgotten about) sets PrivateTmp=yes, the service gets a private /tmp namespace. Your script writes /tmp/agent_log.md, it lands in a directory like /tmp/systemd-private-<hash>-curator.service-<id>/tmp/, and when you cat /tmp/agent_log.md from your own shell afterward, there's nothing there. The file existed. It just existed somewhere you weren't looking, and it gets cleaned up when the service stops.

Then there's aging. /tmp is subject to systemd-tmpfiles cleanup, which I've already been burned by in a different context (see the Plex EAC3 transcoding post). Logs you want to inspect next week shouldn't live somewhere designed to forget things.

Put those together and /tmp is about the worst place to put the output of an unattended agent. The agent's sandbox doesn't consider it in scope, the service manager may give you a private copy, and the OS will clean it up for you.

The fix

The fix has three parts: give the output a permanent home inside a directory Claude is allowed to write, make the calling shell own the log capture, and make denials visible.

Put output in a directory you explicitly grant

#!/usr/bin/env bash
set -euo pipefail

LOG_DIR="$HOME/logs"
mkdir -p "$LOG_DIR"
STAMP="$(date +%Y%m%d-%H%M%S)"
REPORT="$LOG_DIR/sb-curator-out-$STAMP.md"

cd "$HOME/second-brain"   # working dir = what the agent may edit

claude -p "Review the notes in ./inbox. Write a curation report \
to $REPORT listing which notes to promote and why." \
  --add-dir "$LOG_DIR" \
  --allowedTools "Read,Glob,Grep,Write" \
  --permission-mode acceptEdits
Enter fullscreen mode Exit fullscreen mode

The important pieces:

  • cd first. The working directory is the permission boundary. In a systemd unit, the default working directory is / unless you set WorkingDirectory=. Launching an agent with / as its project root is a bad idea for reasons that go well beyond logging.
  • --add-dir extends the boundary to exactly one more directory. That's better than making the boundary bigger.
  • --permission-mode acceptEdits lets file edits inside the granted directories go through without a prompt. Nothing can answer a prompt in headless mode anyway, so the only real question is whether writes inside the boundary are pre-approved or always denied.

To make the grant persistent instead of passing it on every call, put it in the project's .claude/settings.json:

{
  "permissions": {
    "additionalDirectories": ["~/logs"],
    "allow": [
      "Read",
      "Edit(~/logs/**)"
    ]
  }
}
Enter fullscreen mode Exit fullscreen mode

There's a path-syntax trap here too. In Claude Code permission rules, /logs/** does not mean an absolute path. A single leading slash is resolved relative to the settings file. For a real absolute path you need //tmp/** (double slash), or ~/ for your home directory. I've seen people write Edit(/tmp/**), watch it do nothing, and decide the permission system is broken. It works fine, it just reads the path differently than you'd expect.

Let the shell own the log, not the agent

The second change matters more. Stop asking the agent to write its own run log. The agent writes the artifact (the curation report). The shell captures the transcript:

claude -p "$PROMPT" \
  --add-dir "$LOG_DIR" \
  --allowedTools "Read,Glob,Grep,Write" \
  --permission-mode acceptEdits \
  --output-format stream-json --verbose \
  2>&1 | tee "$LOG_DIR/sb-curator-run-$STAMP.jsonl"
Enter fullscreen mode Exit fullscreen mode

A shell redirect or tee runs as your user with normal Unix semantics. Claude's permission boundary doesn't apply to it at all, because Claude isn't the one writing. Now the transcript always exists, even when the agent's own writes get denied, and that's exactly the case where you need it.

The stream-json flag change

That --verbose flag isn't decoration. On current 2.x releases, asking for streaming JSON in print mode without it fails right away:

# Fails: stream-json in print mode now requires --verbose
claude -p "summarize ./inbox" --output-format stream-json
# Error: When using --print, --output-format=stream-json requires --verbose

# Works
claude -p "summarize ./inbox" --output-format stream-json --verbose
Enter fullscreen mode Exit fullscreen mode

This one does fail loudly, so at least it's honest. It still breaks any older script that set up clean machine-readable logs before the requirement existed. If your pipeline stopped producing JSONL after an upgrade, check this first. The research notes I keep point to around v2.1.76 as where it started showing up for people, so check against whatever version you've pinned.

Make denials visible

This is where headless agents are weak on observability. A denied permission isn't an error exit. You have to go looking for it.

With --output-format json or stream-json, the final result message includes a permission_denials array. If it isn't empty, the agent tried something it wasn't allowed to do. That's worth an alert, because it means either your grants are wrong or the agent is going somewhere you didn't expect. Both deserve a look.

RUN_LOG="$LOG_DIR/sb-curator-run-$STAMP.jsonl"

# Last line of stream-json is the result message
DENIALS=$(tail -n 1 "$RUN_LOG" | jq '.permission_denials | length' || true)
DENIALS=${DENIALS:-0}

if [ "$DENIALS" -gt 0 ]; then
  echo "curator: $DENIALS permission denial(s), see $RUN_LOG" >&2
  exit 2   # make systemd mark the unit failed
fi

# Belt and braces: the artifact must exist and be non-empty
[ -s "$REPORT" ] || { echo "curator: report missing" >&2; exit 3; }
Enter fullscreen mode Exit fullscreen mode

Now the timer unit goes red when the agent quietly works around a denial. Before, it just shrugged and exited 0.

The grep exit code detour

Adding checks like these to the curator exposed a second, unrelated bash bug that's common enough to call out. The script counted flagged items in the report:

# Broken: grep -c prints "0" AND exits 1 when there are no matches
FLAGGED=$(grep -ciE 'promote|stale' "$REPORT" || echo 0)
# On zero matches, FLAGGED becomes "0\n0"
# [ "$FLAGGED" -gt 0 ] -> "integer expression expected"
Enter fullscreen mode Exit fullscreen mode

grep -c with no matches prints 0 and exits with status 1. The || echo 0 fallback fires too, so the variable ends up holding two lines. The fix is to swallow the exit code without adding extra output, then default the empty case (for example, when the file doesn't exist):

FLAGGED=$(grep -ciE 'promote|stale' "$REPORT" || true)
FLAGGED=${FLAGGED:-0}
Enter fullscreen mode Exit fullscreen mode

Under set -euo pipefail you need the || true. Otherwise a zero-match grep kills the script at the assignment, which is its own confusing failure.

The systemd side

Last piece: make the unit explicit instead of relying on defaults.

[Service]
Type=oneshot
User=curator
WorkingDirectory=/home/curator/second-brain
ExecStart=/home/curator/bin/curator.sh
# Output lives in $HOME/logs, so PrivateTmp can stay on
PrivateTmp=yes
Enter fullscreen mode Exit fullscreen mode

Once nothing important lives in /tmp, you can leave PrivateTmp=yes on. It's a good hardening default, and it stops being a trap once you aren't depending on /tmp for state.

Why this matters

This is one specific case of a general problem: interactive mode is a different runtime environment than headless mode, and it looks the same only because a human is quietly approving things. Every "yes" you click in a terminal session is an undocumented dependency. The first time the job runs under cron, systemd, CI, or a Kubernetes CronJob, those dependencies disappear at once.

Here's when you'll hit it:

  • Promoting a prompt from "works in my terminal" to "runs on a timer." This is the most common path. Prototyping is interactive, production is headless, and nothing tells you about the gap.
  • Moving scheduled LLM jobs between backends or hosts. When I moved scheduled curation work toward local models, the orchestration moved around too. Every move is another chance for the working directory or temp namespace to change under you.
  • Any agent that writes to shared scratch space. /tmp feels neutral and safe. For an agent it's outside the project boundary, often namespaced, and actively cleaned up.

Here's how I avoid it now:

  1. Every headless agent run gets an explicit working directory and explicit extra directories. No defaults. If I can't list the paths the agent needs to write, I don't understand the job well enough to schedule it.
  2. The agent writes artifacts; the shell writes logs. A transcript captured by tee survives the agent failing, being denied, or improvising. A log the agent writes itself doesn't.
  3. Denials fail the run. permission_denials being non-empty is a non-zero exit. A headless agent that quietly works around a missing permission is worse than one that crashes, because the crash at least shows up in systemctl --failed.
  4. Check that the artifact exists, not just the exit code. [ -s "$REPORT" ] is one line and catches most of the ways this goes wrong.
  5. Nothing that matters lives in /tmp. It's for scratch data you're happy to lose, and "the only record of what my agent did last night" isn't that.

There's also a memory-design angle. In the six-layer memory setup I run for Claude Code, the filesystem layer is the one that assumes the agent can write where it's told to. Headless mode tests that assumption hard. It's the same lesson as promotion pipelines for agent memory: deciding where an agent may write, and proving the write happened, is the hard engineering. The prompt is the easy part.

The sandbox is doing its job here. Scoping an unattended agent's writes to a known set of directories is exactly what you want in production, and I'd much rather have that default than an agent that can scribble anywhere my UID can reach. The failure mode is that the scoping is invisible until you remove the human, and then the only sign is an empty file. Make the boundaries explicit in the command line, capture everything from outside the agent, and treat "the agent adapted around a denial" as the failure it is.

If you're taking agent pipelines from prototype to unattended production and want a second set of eyes on the operational side, that's the kind of work I help teams with.

Top comments (0)