DEV Community

Cover image for Claude might be saturating your machine
Sidhant Panda
Sidhant Panda

Posted on

Claude might be saturating your machine

Identifying orphaned AI tasks via PID 1 checks

My laptop was sitting idle with the fan at full tilt. Nothing was running that I knew of. The culprit turned out to be ten orphaned busy-loops left behind by a Claude Code session two days earlier.

Just fix it for me

If your fan is roaring right now and you would rather not read the rest, paste this into a Claude Code session:

My machine seems idle but the fan is at full speed. Check the load average against my core count, then list the top CPU consumers with their PPID, elapsed time, and full argument list. I am looking for orphaned processes: anything reparented to PID 1 that has been burning CPU for hours or days, especially shells and interpreters where the process name alone tells you nothing. Show me what you find and what each one actually is before killing anything, and let me confirm the list first.

The "show me before killing" part is the important bit, and it is why the prompt is written that way rather than as "find and kill whatever is using CPU." A long-running process pinned to PID 1 is a strong hint that something was abandoned, but it is not proof. Daemons legitimately live there, and your own nohup'd job or a detached build will look identical from a distance. Read the argument list before you agree to anything. In my case the arguments made it obvious in about two seconds.

Finding it

Start with the load average. This is a 10-core machine:

$ uptime
19:39  up 6 days,  6:14, 10 users, load averages: 122.91 167.84 162.08
Enter fullscreen mode Exit fullscreen mode

Load 122 on 10 cores is not idle. Next, the top CPU consumers:

$ ps -Ao pcpu,pid,ppid,user,comm -r | head -12
 %CPU   PID  PPID USER   COMM
139.8  8320     1 user   /Applications/Google Chrome.app/Contents/MacOS/Google Chrome
 60.9 94281     1 user   /bin/zsh
 59.4 94279     1 user   /bin/zsh
 59.4 94288     1 user   /bin/zsh
 59.1 94287     1 user   /bin/zsh
 57.9 94285     1 user   /bin/zsh
 57.6 94284     1 user   /bin/zsh
 57.6 94286     1 user   /bin/zsh
 57.3 94283     1 user   /bin/zsh
 55.9 94282     1 user   /bin/zsh
 51.1 94280     1 user   /bin/zsh
Enter fullscreen mode Exit fullscreen mode

Ten zsh processes at roughly 60% each, and every one has PPID 1. That is the tell: their parent died and launchd adopted them. comm only gives you the binary name, so pull the full argument list to see what they actually are:

$ ps -o pid,lstart,etime,pcpu,args -p 94279,94280,94281
Enter fullscreen mode Exit fullscreen mode

Note the comma-separated list. Passing -p alongside -A silently gets you every process on the box instead.

The arguments explained everything:

/bin/zsh -c source ~/.claude/shell-snapshots/snapshot-zsh-XXXX.sh 2>/dev/null || true && eval '
SP=/private/tmp/claude-501/<project>/<session-id>/scratchpad
# saturate all cores, then run the suite under contention
NCPU=$(sysctl -n hw.ncpu)
for i in $(seq 1 $NCPU); do (while :; do :; done) & done
LOADPIDS=$(jobs -p)
pnpm test:integration > "$SP/load.log" 2>&1
kill $LOADPIDS 2>/dev/null
...'
Enter fullscreen mode Exit fullscreen mode

Elapsed time was 01-22:41:17. Just under two days of while :; do :; done on every core.

A previous session had deliberately pegged all ten cores to run an integration suite under CPU contention, then meant to clean up after itself. The test run itself was long gone, confirmed by the absence of any vitest or pnpm process:

$ ps -Ao pid,ppid,pcpu,etime,comm | grep -Ei "[n]ode|[v]itest|[p]npm"
58938 58933   0.0    00:59 .../bin/node
88656 88648   0.0 02-00:57:58 .../bin/node
Enter fullscreen mode Exit fullscreen mode

Only the spinners were left. They accounted for about 700% of the 850% total CPU in use:

$ ps -Ao pid,ppid,pcpu,comm | awk 'NR>1 && $3>20 {sum+=$3; n++} END {print "procs >20% CPU:", n, " total %CPU:", sum}'
procs >20% CPU: 12  total %CPU: 850.3
Enter fullscreen mode Exit fullscreen mode

Why the cleanup failed

Two things went wrong, and either alone would have been enough.

LOADPIDS=$(jobs -p) was the first. Job control is off in a non-interactive shell, so jobs -p returned nothing and the cleanup kill had no arguments to act on. Collect $! after each background spawn instead:

LOADPIDS=""
for i in $(seq 1 $NCPU); do
  (while :; do :; done) &
  LOADPIDS="$LOADPIDS $!"
done
Enter fullscreen mode Exit fullscreen mode

The second is that the parent shell died before reaching the kill line at all. A cleanup step on the happy path is not cleanup. Use a trap so it fires even on interrupt:

trap 'kill $LOADPIDS 2>/dev/null' EXIT INT TERM
Enter fullscreen mode Exit fullscreen mode

Fixing it

Plain kill was enough. No -9 needed:

$ kill 94279 94280 94281 94282 94283 94284 94285 94286 94287 94288
$ ps -o pid= -p 94279,94280,94281,94282,94283,94284,94285,94286,94287,94288 | wc -l
0
Enter fullscreen mode Exit fullscreen mode

Afterwards, Chrome was back on top where it belongs and nothing else broke 30%:

$ ps -Ao pcpu,pid,comm -r | head -4
 %CPU   PID  COMM
129.6  8320  /Applications/Google Chrome.app/Contents/MacOS/Google Chrome
 43.2 58249  .../Google Chrome Helper (Renderer)
 27.5 69725  .../UsageTrackingAgent
Enter fullscreen mode Exit fullscreen mode

The load average still read 74 at that point, which is expected. It is a decaying rolling average and lags reality by design. The 1-minute figure had already dropped from 122 to 74 and kept falling; the 15-minute figure was still carrying the previous quarter-hour of saturation. Judge the fix by the process list, not the load average. The fan takes another minute or two beyond the CPU drop while the heat already in the chassis dissipates.

Checking your own machine

uptime                              # load average well above your core count?
sysctl -n hw.ncpu                   # what your core count actually is
ps -Ao pcpu,pid,ppid,user,comm -r | head -15
Enter fullscreen mode Exit fullscreen mode

Look for processes eating CPU with PPID 1 and a long etime. Reparenting to launchd plus an elapsed time measured in days means nobody is supervising it and nobody is going to clean it up. Shells and interpreters are worth a closer look, since comm shows you /bin/zsh or node and tells you nothing about what is inside. Get the arguments:

ps -o pid,ppid,lstart,etime,pcpu,args -p <pid>
Enter fullscreen mode Exit fullscreen mode

On macOS, Activity Monitor sorted by CPU shows the same thing, but the process name column has the same problem: ten identical zsh rows and no hint of what they are running.

Takeaway

Agents run real commands, including ones that deliberately consume resources, and a crashed or cancelled session does not necessarily take its children with it. If you let an agent spawn background work, it is worth knowing how to spot the leftovers. ps -Ao pcpu,pid,ppid,user,comm -r | head costs nothing and would have caught this two days earlier.

Top comments (7)

Collapse
 
motedb profile image
mote

I hit this exact thing last month — six orphaned Node processes from a Cursor session, each pinning a core at 100%. The PPID=1 check is the key signal. A lot of people grep for high CPU and miss that the parent is init, which is what tells you "this was abandoned" versus "this is just a busy daemon."

One shortcut I've started using: pgrep -P 1 -l filters to just reparented processes without scanning the full ps table. Pair it with top -l 1 -o cpu -n 10 on macOS and you get a quick two-step triage.

The "show me before killing" framing in your prompt is better than what I was doing — I killed a launchd daemon once that looked suspicious and spent the next hour figuring out why Bluetooth stopped working. Have you considered wrapping this into a shell alias that runs on Claude Code session end?

Collapse
 
vinimabreu profile image
Vinicius Pereira

"jobs -p returning nothing in a non-interactive shell is such a quiet way to lose" is the perfect diagnosis, and the trap fix is right. One thing worth naming, since your own conclusion points at it: a crashed or cancelled session does not necessarily take its children with it, and a trap does not fire on SIGKILL or when the parent is killed hard from outside. So the trap covers the clean-exit and Ctrl-C paths, which is most of them, but not the exact case your title is about, the session that died abnormally.

The belt-and-suspenders for agent-spawned workloads is to make the children unable to outlive the parent regardless of how it dies. On Linux, run them in their own process group and have a supervisor send the group a signal, or use a PR_SET_PDEATHSIG so the kernel kills the child when the parent exits. On macOS there is no pdeathsig, so the practical equivalent is spawning under a wrapper that polls getppid() and exits when it becomes 1. The PPID 1 you used as the diagnostic tell is the same signal the child can watch to self-terminate. Cleanup you own can be skipped; a death condition the OS or the child enforces cannot.

Collapse
 
eduzsh profile image
Edu Peralta

I have hit a version of this exact failure mode with background shells that agents spawn. The part that stood out to me is that both failure paths you found, jobs -p being empty in a non-interactive shell and the parent dying before its own cleanup line, are invisible from inside the session itself. The agent thinks it finished cleanly because its last command exited zero, and nothing about that exit code tells you a grandchild process is still out there in a while loop. I have started treating trap based cleanup as mandatory for anything running in the background, not optional, since the whole point of background work is that nobody is watching it end.

Collapse
 
nazar-boyko profile image
Nazar Boyko

I can be wrong but jobs -p returning nothing in a non-interactive shell is such a quiet way to lose, because the cleanup line is right there in the script and reads like it works. The trap is the real fix, and I'd argue it's the fix even when you collect $! correctly, since the parent dying is the more common failure of the two. Worth noting the trap won't save you from a kill -9 on the parent either, so for anything that pegs every core I'd have the loop time-limit itself as a backstop. Two days of while :; do :; done is a great argument that the spinner should never outlive its purpose regardless of who's supposed to clean it up.

Collapse
 
mateo_ruiz_6992b1fce47843 profile image
Mateo Ruiz

This is a good example of why agent reliability extends beyond prompt quality. Once an agent can spawn subprocesses, you've crossed into systems engineering territory. Every background task should have a lifecycle: ownership, timeout, heartbeat, and guaranteed cleanup even if the parent crashes or the session is interrupted. We've seen similar issues when agents launch long-running builds, local inference servers, or stress tests. A trap helps, but process groups, watchdogs, and explicit resource budgets make these failures much harder to trigger. AI agents don't just generate code anymore they're operating systems clients, and they need to be treated with the same discipline as any other automation.

Collapse
 
publiflow profile image
PubliFlow

Good technical content. I'd love to see more exploration of the edge cases and failure modes — understanding when and why these patterns break down is often more valuable than knowing how to implement them.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.