I set up alerts for my servers. I write watchdog jobs. I go to real trouble to make sure I hear about it when a service misses two consecutive checks. This morning a thought arrived: who is watching the machine I write all of that on?
So I sat down and counted. launchctl list told me there are 561 jobs in my own session. Then I opened that number up and found that the figure which frightened me said nothing at all, while the one that was genuinely broken never said a word.
One detail worth recording: the first count said 554, and seven minutes later it said 561. The list moves while you count it. Every number below is derived from a single snapshot taken at 08:41:38, otherwise the totals wouldn't add up.
192 of My Jobs Were Killed, and None of Them Failed
My first reading of the table went like this: of 561 jobs, 201 have a last exit status other than zero. Thirty-six percent. The kind of ratio that makes your hands go cold.
Then I noticed that 192 of those 201 were -9. Note the minus sign. man launchctl is perfectly clear at exactly this point: if the number in the second column is negative, it represents the negative of the signal that killed the job. So -9 isn't an error code, it's SIGKILL.
Here I tried to draw a conclusion, and I was wrong. I was about to write "all 192 signal-terminated jobs are SIGKILL, not a single SIGTERM" and make something of it. The EnableTransactions entry in man launchd.plist stopped me:
When launchd stops an active process, it sends SIGTERM first, and then SIGKILL after a reasonable timeout. If the process is inactive, SIGKILL is sent immediately.
So the polite path exists and works — for active processes. The jobs I caught sitting idle in a snapshot never enter that path by definition. 192 out of 192 SIGKILL isn't a scandal, it's a point-for-point confirmation of documented behaviour. And launchctl list only shows the last exit anyway; a job that handled SIGTERM properly and exited zero is invisible in that column by construction.
Separating out the negatives collapsed the table to this:
| Last exit status | Count |
|---|---|
| Clean exit (0) | 360 |
| Terminated by signal (SIGKILL) | 192 |
| Real error (> 0) | 9 |
| Total | 561 |
Nine. What I had read as thirty-six percent was one and a half. The frightening number came from me reading the wrong column.
Eight of the Nine Errors Are Mine
I printed all nine. This is where the genuinely embarrassing part begins:
| Exit code | Job |
|---|---|
| 127 | app.meos.model |
| 255 | com.secwise.panel-tunnel |
| 78 | com.openai.atlas.update-helper |
| 1 | app.meos.console |
| 1 | app.meos.monitor |
| 1 | com.erbay.spark-bulut |
| 1 | homebrew.mxcl.postgresql@16 |
| 1 | local.sahinkulhelpdesk.hdpreview-observer |
| 1 | com.apple.Siri.agent |
In that same snapshot, 511 of the 561 jobs start with com.apple.*. The plist files on disk show the same ratio: 887 of them on Apple's sealed system volume, a grand total of 47 installed afterwards.
But only one of the nine broken jobs is Apple's. Eight are either ones I wrote or ones belonging to tools I installed. I went looking for things that had quietly broken on my own machine, and I found myself on the list.
Two of those codes are talkative. 127 is /bin/sh saying "there is no such command." And 78 is EX_CONFIG in Apple's own sysexits.h header — a configuration error.
exit 127: The Script Is Gone, the Job Is Still Trying
I looked at the plist for app.meos.model. It invokes a shell script through /bin/sh: deploy/finetune/model-serve.sh. I checked that path.
No directory. No script. Not even the parent folder.
The tail of the error log says the same thing, many times over:
/bin/sh: .../deploy/finetune/model-serve.sh: No such file or directory
I honestly cannot tell you how long the job has been broken: that project isn't a git repository, so there's no record of the day the script was deleted. What I have are two lower bounds. The plist has been sitting there since 5 June, and the error log file was created on 13 September and was still being written to at 08:36 this morning.
The second broken job is a sibling in the same family. app.meos.console invokes a Python command; the package is gone, and what's left behind is ModuleNotFoundError: No module named 'meos.cli'.
Both of those have KeepAlive enabled in their plist. In other words, I told launchd "if this one falls over, pick it back up." And launchd is keeping its word.
The third, app.meos.monitor, throws the same Python error but runs on an entirely different mechanism: its plist has no KeepAlive, it has StartInterval 900. That isn't a loop, it's a timer that wakes up every fifteen minutes. Same error, two different tempos — and that is exactly why one of the log files you're about to see is fifty times smaller than another. Of the three dead jobs, two are loops and one is a timer.
I Counted for Two Minutes: 20.0 and 10.9 Seconds
Guessing whether the loop was still running wasn't good enough, so it went on a stopwatch. Line counts of the two error logs, then 120 seconds of waiting, then the counts again.
In those two minutes app.meos.model tried 6 times: once every 20.0 seconds. The ThrottleInterval value I typed into that plist by hand is exactly 20.
app.meos.console tried 11 times: once every 10.9 seconds. That job's plist has no ThrottleInterval. Apple's launchd.plist manual says what happens in that case: by default, jobs will not be spawned more than once every 10 seconds.
Three independent measurement paths all landed in the same place:
| Measurement path | model |
console |
|---|---|---|
| Live 120-second sample | 20.0 s | 10.9 s |
| launchd counter ÷ 48.7 h uptime | 23.9 s | 12.1 s |
| Log file lines ÷ file age | 26.1 s | 12.2 s |
| Floor the configuration specifies | 20 s | 10 s |
The long-window averages sitting slightly above the floor is the expected outcome — the machine sleeps in between. The live sample sits right on the floor.
There are two ways to state the daily total, and the honest thing is to give both. Multiply the live sample out to twenty-four hours and you get 12,240 failed starts; that's an upper bound, because it assumes the machine never sleeps. Divide launchd's own counter by uptime and you get what actually happened: roughly 10,800 per day. Take whichever you like — both are close to five digits. And none of them break anything. They just never stop.
25 MiB, Five Sentences
I counted the three error logs. The values below are a snapshot at 08:48:43:
| File | Lines | Unique lines | Size |
|---|---|---|---|
model.err.log |
57,569 | 1 | 5.8 MiB |
console.err.log |
397,840 | 4 | 19.3 MiB |
launchd.err.log |
8,220 | 4 | 0.4 MiB |
The first two are growing while you read this line. The third, belonging to monitor the timer, adds four lines every fifteen minutes — that's why it's fifty times smaller.
Across the three files there are four hundred and sixty-three thousand lines. But the unique line count isn't nine, it's five: the two Python jobs write the byte-for-byte identical four-line traceback, and counting those four twice would be cheating. On top of that sits model's single-line message. A total of 25.4 MiB of disk — add the column yourself and rounding will show you 25.5 — all of it copies of five sentences.
Nothing rotates a file written to a path of my own choosing through StandardErrorPath — not newsyslog, not anything else. Multiply that daily growth rate out to a year and you get roughly 625 MiB per year of identical error text — a derived number, not a measured one, but the tempo behind it is measured. For three of them. For three jobs that do nothing.
On the days I went hunting for disk space, I never once looked in that folder. It wasn't in the accounting where I promised my laptop 60 GiB and found 14 either.
It Wasn't Hidden. It Was Drowned.
Up to this point the story in my head was: macOS hid this from me. I went to the system log, ready to confirm my hypothesis.
The query returned zero lines. Not just for my jobs — zero for everything. The entire system log for the last two minutes: zero lines.
If the scanner is broken, there is no finding. The reason was stupid: my shell has a builtin called log that shadows /usr/bin/log. I called it with the full path.
Last two minutes: 102,028 lines. And my hypothesis collapsed.
In the last thirty minutes, the launchd process log contains 5,694 lines mentioning meos. In the words of man log, log show displays "contents of the system log datastore" — so those lines were written somewhere, and I was able to query them half an hour later. They also arrived without me asking for --info or --debug.
And launchd says what it's doing, word for word:
[gui/501/app.meos.console:] service state: spawning
[gui/501/app.meos.console:] launching: inefficient
Read that second line. launchd labelled this launch "inefficient" itself. In that same thirty minutes there are 507 launching: inefficient lines across the whole system, and 268 of them — 52.9 percent — come from these jobs. More than half of the inefficient-launch warnings on my machine belong to two loops whose programs don't exist.
So why did I never see it? In that same thirty-minute window the system log wrote 1,336,483 lines in total. Forty-four thousand lines a minute. My broken jobs make up 0.426 percent of that: one line in 235.
It wasn't hidden. It was drowned. The difference between those two is the difference in who is at fault. Writing about alert fatigue from the outside was easy; running into a 1-in-235 ratio on your own machine teaches it differently.
launchd Had Already Counted, and I Had Already Looked
I saved the most irritating part for last. There is a subcommand called launchctl print, and it dumps everything it knows about a single job. The relevant lines from the output, reordered:
state = spawn scheduled
runs = 7354
last exit code = 127
minimum runtime = 20
runs = 7354. launchd counted the attempts. For the other job, runs = 14527. Together, 21,881 attempts within 48.7 hours of uptime.
I was going to end this section with "I never typed that command." I can't, because it would be a lie.
I typed that command on 15 September. The piece I wrote about ssh -J contains a section titled "12,342 attempts": same Mac, same launchctl print, same runs counter, same ThrottleInterval source. Sixteen days ago I learned what this tool does and wrote it up.
And then I never pointed it at anything else. In fact the agent that article was about — com.secwise.panel-tunnel — is sitting on the second row of the nine-error table above, still returning 255. Learning a tool, answering exactly one question with it, and stopping there is more uncomfortable than never learning it. In the second case you don't know; in the first you know and you don't ask.
One more note for honesty's sake: man launchctl says in capital letters that the print output is "NOT API in any sense at all," and warns you not to rely on its structure. That sentence is in the man page shipping with macOS 26.6.2; the version in Apple's open source repository predates launchctl 2.0 and contains neither print nor bootout. So launchd hands you the number but refuses to stand behind it. It isn't a surface a watchdog script can parse safely.
Being Registered and Working Are Not the Same Question
macOS keeps a ledger of third-party background items. You read it with sfltool dumpbtm — and on this machine, on macOS 26.6.2, I read it without root, as ordinary user 501.
The ledger holds 126 records. Forty-six are legacy-style agents and daemons — near-identical to the 47 plists I counted on disk. Twenty records list their developer as "Unknown Developer."
My three dead jobs are in the ledger too. Their names, their types, which plist they came from — it's all there.
That they are broken is not.
The ledger answers the existence question: what is registered on this machine? There is no line answering the health question. Apple's modern path, SMAppService (macOS 13 and later), exists so that applications can register helpers that live inside their own bundle — not for a shell script sitting in a folder of mine. A LaunchAgent I wrote by hand still runs and is still documented; but the name it goes by in Apple's own documentation is now "legacy plist." Not deprecated, just out of favour. And unsupervised either way. That gap points at me, not at Apple.
What I'd Check on Any Other Mac
After this morning, my order of operations is:
- In the
launchctl listoutput, filter for positive exit codes. The negatives are signals; they're noise. (The man page says oflist: "Recommended alternative subcommand: print."listis still the practical front door that shows every job on one screen, but switch toprintwhen you go deeper.) - Of the positives, look at the non-
com.apple.*ones first. The thing you installed yourself is the thing most dependent on your maintenance. - Verify that the program path in each plist actually exists on disk.
exit 127is telling you this, but only if you ask. - Check the
runscounter withlaunchctl print. If you see a four-digit number, that isn't a service, it's a loop. - Look at the size and the unique line count of the files you write through
StandardErrorPath. If the unique count is close to one, there's no information in there — only volume. - When you stop using a job, remove it — but mind three things.
Item six needs unpacking, because it's the only real lesson in this piece and leaving it half-stated is useless.
You stop your own job with launchctl bootout gui/$UID/<label>. But bootout alone isn't permanent: as long as the plist stays in ~/Library/LaunchAgents, the job comes back at your next login. If you want it gone, delete the plist too; if you deliberately want to keep it around, disable it with launchctl disable gui/$UID/<label>.
Second: don't touch the 887 files in /System/Library/LaunchAgents and /System/Library/LaunchDaemons. They sit on the sealed system volume protected by SIP; trying to tamper with them either fails or breaks your system. Booting out com.apple.* jobs in your own gui domain is a bad idea too — you're dismantling working parts of your session.
Third: for third-party apps' background items you may not need the terminal at all. The Login Items pane in System Settings is the user-facing face of those 126 records sfltool dumpbtm prints. If you want to switch off an app's updater, that's the right place.
Deleting the project does not delete the job. Every one of the 25 MiB in this piece is the price of skipping that sentence.
Closing
If a service on my servers misses two checks, my phone rings. What I learned this morning isn't that I never extended that same care to the machine I work on — it's that the care I did extend was aimed at the wrong question.
Every monitoring mechanism I've built asks "is it working?" None of them asks "is it still there?" A job whose program has been deleted falls outside the scope of the first question: it doesn't crash, it doesn't slow down, it doesn't return a wrong answer. It just writes its own absence ten-thousand-odd times a day into a file nobody reads.
The fault that goes unnoticed longest in a system isn't the one that breaks something; it's the one that never manages to do anything at all. Our dashboards are built to watch what runs. What never ran, you only see when you decide to sit down and count — and for that counting, having used the tool once is not enough.
Official Sources
-
launchd.plist(5) — Apple open source distributions —
ThrottleIntervaland the default 10-second spawn limit -
launchctl(1) — Apple open source distributions — a negative value in
listoutput meaning a signal -
sysexits.h — Apple Libc —
EX_CONFIGbeing exit code 78 - SMAppService — Apple Developer — registering in-bundle helpers on macOS 13 and later
- Change login items on Mac — Apple Support — the login items surface shown to the user
The SIGTERM/SIGKILL ordering under EnableTransactions, the "NOT API" warning on print output, the recommendation of print over list, and the datastore definition of log show are taken from the man launchd.plist, man launchctl and man log pages shipping with macOS 26.6.2. Those items do not appear in the open source versions linked above.
Top comments (0)