Most small IT teams don't have a tooling problem. They have a nobody-was-assigned-it problem.
The fix isn't a platform. It's five small cron jobs, each one deleting a recurring manual check forever. Here they are, with the actual commands.
1. Disk-full early warning
The classic 3am outage: /var hits 100% and nothing logs anymore. Warn at 80%, don't wait for the fire.
#!/usr/bin/env bash
# /opt/scripts/disk_alert.sh — cron: */30 * * * *
ALERT=80
df -H --output=pcent,target | tail -n +2 | while read -r pct mount; do
used=${pct//%/}
[ "$used" -ge "$ALERT" ] && echo "DISK ${mount} at ${pct}" \
| mail -s "disk alert $(hostname)" ops@yourco.example
done
Run it every 30 minutes. One afternoon of work buys you the difference between a note in your inbox and a dead server.
2. Backups that verify themselves
A backup you've never restored is a rumour. Schedule the proof:
# cron: 0 4 * * 1 — verify newest backup every Monday
LATEST=$(ls -t /backups/*.tar.gz | head -1)
tar -tzf "$LATEST" > /dev/null && echo "OK $LATEST" >> /var/log/backup_verify.log \
|| echo "FAIL $LATEST" | mail -s "BACKUP VERIFY FAILED" ops@yourco.example
Archive integrity isn't a restore test, but it catches the silent-corruption case that kills most "we have backups" stories.
3. Certificate expiry countdown
Let's Encrypt renewals fail quietly — you find out from users, after the browser warning. Check daily from the outside:
# cron: 0 6 * * *
for d in app.yourco.example portal.yourco.example; do
days=$(( ($(date -d "$(echo | openssl s_client -servername "$d" -connect "$d:443" 2>/dev/null | openssl x509 -noout -enddate | cut -d= -f2)" +%s) - $(date +%s) ) / 86400 ))
[ "$days" -lt 14 ] && echo "$d expires in $days days" \
| mail -s "cert expiring: $d" ops@yourco.example
done
4. Failed services, named and shamed
systemd restarts most things — but which services needed rescuing this week tells you what's actually fragile:
# cron: 0 7 * * 1 — weekly fragility report
journalctl --since "7 days ago" -p err -u '*' --no-pager \
| grep -iE "failed|restart" | sort | uniq -c | sort -rn | head -20 \
| mail -s "weekly service fragility report" ops@yourco.example
Read this list in your Monday coffee. Whatever tops it is next sprint's real work.
5. Weekly change inventory
"Who changed what?" after an incident is answered in seconds when a diff lands in your inbox automatically:
# cron: 0 18 * * 5 — Friday change report
etckeeper commit "weekly auto" 2>/dev/null || true
diff /var/lib/etckeeper-last.conf /etc/etckeeper/etckeeper.conf >/dev/null 2>&1 || true
Even a crude /etc snapshot emailed weekly beats archaeology with stat at 2am.
None of this is clever. That's the point — these are the checks every team reinvents badly at 3am, once. Ship them once, properly, and they run quietly for years.
If you'd rather not type all five yourself, we've packaged these as a ready-to-run Automation Starter Pack (all scripts, sane defaults, install notes):
👉 https://hive80lab.gumroad.com/l/automation-starter-pack
And if incident nights are your recurring nightmare, the Ops Starter Kit has the comms templates, backout decision points and post-incident review formats that pair with the monitoring above:
👉 https://hive80lab.gumroad.com/l/ops-starter-kit
Ship one today. Ship the rest this week. Your on-call self will notice.
Top comments (0)