The outage nobody schedules
Ask anyone who runs a handful of servers: the most common 3am incident is still No space left on device. Not because disks are mysterious — because disk-full is a slow failure that only becomes loud at the worst moment: a log file grows 40GB, a build cache never rotates, a database temp dir spikes, and suddenly the app that ran fine for a year won't start.
Monitoring dashboards notice. Your phone at 3am doesn't. Here's the watchdog that closes the gap.
The 15-line watchdog
Run this from cron every 10 minutes:
#!/bin/bash
PCT=$(df / | awk 'NR==2 {gsub("%",""); print $5}')
LIMIT=85
if [ "$PCT" -ge "$LIMIT" ]; then
curl -s "https://your-alert-hook/?disk=${PCT}&host=$(hostname)" > /dev/null
logger "DISKWATCH: ${PCT}% used on / - alert sent"
fi
Three details that make it actually work:
- Alert on a webhook, not email. Email from cron dies quietly in a spam folder. A webhook that hits your phone (Slack, Telegram, pushover-style) is the only alert that reliably wakes a human.
- Escalate, don't just alert once. If disk is still above the limit on the next run, the alert re-fires — that's the point of running it from cron. Silence after a missed alert is the killer.
-
Log one line locally.
loggerwrites to the system log, so even if the network is the thing that's down, you have a forensic trail of when the disk crossed the line.
Then find the actual eater
Disk-full is a symptom; the disease is a process nobody taught to rotate. The 30-second triage:
du -xh / --max-depth=2 2>/dev/null | sort -rh | head -20
Nine times out of ten it's /var/log, a build cache, or an app writing backups next to itself instead of to object storage. Fix the writer, not the fullness.
Make the pattern permanent
One watchdog on one box is a good night's sleep. The next level is a small library of these guards — disk, cert expiry, backup freshness, hung jobs — deployed identically everywhere, with a runbook per failure mode so whoever gets paged knows the fix without reading the code.
I package exactly that:
- Ops Starter Kit — https://hive80lab.gumroad.com/l/ops-starter-kit
- Automation Starter Pack — https://hive80lab.gumroad.com/l/automation-starter-pack
- Agent-Ops 24/7 — https://hive80lab.gumroad.com/l/agent-ops-24-7
Either way: put the watchdog on one box today. 85% limit, webhook alert, cron every 10 minutes. That's fifteen lines standing between you and the most predictable outage in IT.
Boring automation for small IT teams. More at https://hive80lab.gumroad.com
Top comments (0)