DEV Community

Sam Hartley
Sam Hartley

Posted on

My Pipeline Went Silent for Five Hours — Every Job Failed the Same Quiet Way

My Pipeline Went Silent for Five Hours — Every Job Failed the Same Quiet Way

I run a small scanner pipeline on a Mac in a closet. A handful of Python scripts on schedules, all reading and writing to one project folder, a couple of them talking to an exchange API. It's the kind of setup that hums along for months and quietly convinces you it has no moving parts.

Then one morning it did nothing for five hours — and the first I heard about it was me, sitting down at my desk, noticing that a normally chatty system had posted nothing since before dawn.

No error. No alert. No red anything. Just five hours of a pipeline that was completely, quietly dead.

This is the postmortem of a single point of failure I'd built without noticing, and the two changes that mean it can't fail this silently again.

The clue was an absence

The honest version of how I noticed: I didn't. Nothing paged me. I only realized because the volume of output I expect by mid-morning wasn't there. When your monitoring is "I check it and it looks busy," a dead system and a calm system are the same picture.

So I went looking for errors and found none — which is exactly the problem. The jobs weren't crashing. They were failing in the one way that never generates a complaint: they never got to the point of running.

The thing everything depended on

Every scheduled job had one thing in common: its working directory lived on a network share mounted at /Volumes/share.

mount | grep share
# //user@nas/… on /Volumes/share (smbfs, nodev, nosuid, …)
Enter fullscreen mode Exit fullscreen mode

The code was fine. The schedules were fine. But the entire pipeline's ability to do anything at all rested on that mount being present. And last night it wasn't.

The timestamp that started everything:

02:35:43  [TRANSITION] OK -> UNMOUNTED
Enter fullscreen mode Exit fullscreen mode

From 02:35 on, a mount point that every script treated as a normal directory was simply not there.

Anatomy of a silent failure

Here's what "the share is gone" does to a set of scripts that assume it's there — and why none of it screams:

  • Scheduled jobs fail at startup. A cron or launchd job whose script lives on the missing share can't even begin. If its output goes nowhere anyone reads, the failure is invisible.
  • Cron mail is off. The classic line MAILTO="" in the crontab — I set it years ago to stop noise — means a job that dies produces no message at all. A silent crontab plus a missing dependency equals nothing in your inbox, forever.
  • Processes that do start hang on I/O. My tracker process was alive but wedged:
  TimeoutError: [Errno 60] Operation timed out
    ... /Volumes/share/shadow_data/correlations.jsonl
Enter fullscreen mode Exit fullscreen mode

A process that's stuck isn't a process that's dead — to any "is the PID running?" check, it looks healthy.

The paper trail was just a hole in a log. My fill-watch log's last line before the outage was at 02:34. The next line was at 07:34, after I fixed it by hand. Five hours with no entries, and nothing that treated an empty log as an event.

The bug that made it worse

If the mount had simply failed and retried, this might have been a ten-minute blip. But my auto-mount helper turned a blip into a five-hour outage.

I had a launchd job that ran every five minutes and tried to remount the share. It worked like this:

# auto-mount, every 5 min
mount_share() {
  mkdir -p "$MOUNT_DIR"
  osascript <<'EOF'
    mount volume "smb://user@nas/share"
  EOF
}
Enter fullscreen mode Exit fullscreen mode

That osascript call has no timeout. Last night it hung — and stayed hung for five hours. Worse, the script held a lock directory while it waited, so the next run five minutes later found the lock held and bailed out. The self-healing loop had blocked the very recovery it existed to perform.

02:39:48  mount attempt starts  → hangs
…          lock held, every later attempt exits early
07:xx     I kill it by hand, clear the lock, mount manually
Enter fullscreen mode Exit fullscreen mode

This is the part that stung: the thing I'd built to fix the outage was the reason it lasted all morning.

The fix, part one: every network call gets a deadline

The remount logic now wraps the blocking call in a hard timeout:

run_with_timeout() {          # default 25s, override via SMB_MOUNT_TIMEOUT_SECONDS
  local secs="${SMB_MOUNT_TIMEOUT_SECONDS:-25}"
  ( "$@" ) & local pid=$!
  ( sleep "$secs"; kill -9 $pid 2>/dev/null ) & local watcher=$!
  wait $pid; local rc=$?
  kill $watcher 2>/dev/null
  return $rc
}

mount_share() {
  mkdir -p "$MOUNT_DIR"
  run_with_timeout osascript -e "mount volume \"smb://user@nas/share\""
}
Enter fullscreen mode Exit fullscreen mode

I tested it against a deliberately hanging mount:

hanging osascript → rc=124 after 4s   (previously: hung for 5 hours)
Enter fullscreen mode Exit fullscreen mode

A hung mount now dies in seconds, releases the lock, and the next cycle can try again. The rule I should have applied years ago: any call that touches the network gets a deadline. Retries without timeouts aren't resilience — they're a new lock waiting to happen.

The fix, part two: stop depending on the share

The deeper fix was to remove the single point of failure rather than babysit it. Two changes:

  1. Jobs run from local disk. The working directory moved off the network share entirely. The mount can now be missing and the scripts still run — they just can't reach their data, which is a state they're written to detect and report.
  2. Logging is local-first, and shipped to the NAS as a delta. Writers always append to a local file, so the network can never block them. A small shipper process pushes only the new bytes since its last offset:
   local log → offset state → append delta to NAS copy
Enter fullscreen mode Exit fullscreen mode

If the NAS is down, the shipper buffers locally and catches up when it returns — no lost lines, no duplicates. The canonical copy still lives on the NAS; the local copy is the buffer that keeps the pipeline from caring whether the NAS is up.

Now a dead NAS degrades the pipeline instead of killing it — and, crucially, the pipeline keeps producing the evidence I need to see that something's wrong.

The fix, part three: a watchdog that alerts on absence

Last piece: something that watches for the absence of a healthy share, not just the presence of errors.

A tiny watcher runs every five minutes and only speaks on a state change:

  • Share down for more than 15 minutes → one alert. Not every five minutes, once.
  • Share comes back → a recovery message, including how many bytes the shipper had buffered while it was away.

It's written fail-safe — it swallows its own exceptions so a broken watchdog can't take down what it's watching — and it stays quiet under the threshold, because an alert you learn to ignore is worse than no alert.

What I'd tell myself before the next outage

  1. "Nothing is broken" and "nothing ran" look identical. If your monitoring only fires on errors, a dependency that's simply gone produces total silence — and silence reads as success. Monitor for the absence of expected effects, not just the presence of failures.

  2. A self-healing loop without a timeout is worse than none. Mine held a lock for five hours and blocked its own recovery. Every retry, mount, and network call needs a hard deadline, or your fix becomes the outage.

  3. If every job shares one dependency, its failure mode is "everything looks fine." And the monitor must never live behind the thing it monitors — mine did, on the same share, which is why the outage stayed invisible.

The cost of this one was an annoying morning. The lesson is cheaper than the next one: on a shared dependency, silence isn't health. It's just the shape of a failure you haven't built a question for yet.


Has anyone else been burned by a "self-healing" script that needed healing itself — or by monitoring that couldn't see an absence? I'd genuinely like to know how people detect "the scheduled thing didn't run" without inventing a whole observability stack for three Python scripts. Drop your approach in the comments.

Part of my Building in Public series — previously: the scheduled job that died with 'Operation not permitted', the exit rule that counted the wrong thing, and the drawdown breaker watching a number that never moved.

Top comments (0)