Originally published on kuryzhev.cloud
You wrote a bash backup script, ran it by hand, and watched the files land in S3. A week later someone asks for a restore and the newest object in the bucket is nine days old. This is a typical failure: the bash backup script works in your terminal and quietly does nothing, or something different, when cron runs it. This runbook walks from symptom to cause to three fixes.
Symptoms
Match what you see against this list before changing anything.
- The script succeeds interactively, but the S3 bucket has no new objects after the scheduled time.
- The cron log shows the job started, yet no backup log is written. Where to look depends on the host:
-
grep CRON /var/log/syslogon Debian or Ubuntu with rsyslog. -
/var/log/cronon RHEL-family systems. -
journalctl -u cronorjournalctl -u crondwhere only the journal exists.
-
- Errors such as
aws: command not found,Unable to locate credentials, orThe config profile (...) could not be foundappear only in local mail or not at all. - The backup log grows without limit, or the disk fills with archives in
/var/backups. - Two backup processes show up in
psat once, and the second one is still running when the next schedule fires. - The log says "done" but the S3 prefix is missing files, or files deleted at the source are still in the bucket.
If the cron daemon never logged the job at all, the problem is the crontab itself. Common causes:
- The job is in the wrong user's crontab.
- The schedule line has a syntax error.
- A file in
/etc/cron.dhas a dot in its name, is writable by group or other, or lacks a trailing newline.
Check that first with crontab -l as the intended user, or inspect the /etc/cron.d file.
Root cause: cron is not your shell
Cron runs jobs with a minimal environment. The common default on Linux cron implementations is:
- a short
PATH(often/usr/bin:/bin); -
SHELL=/bin/sh; - no TTY;
- none of the exports from your
~/.bashrcor profile.
Anything your interactive session provided silently is missing. That includes a pip-installed or /usr/local/bin AWS CLI, AWS_PROFILE, SSO sessions, relative paths, and even the working directory.
Three further causes sit on top of that, and they explain the "ran but wrong" symptoms:
-
Swallowed exit codes. Without
set -euo pipefail, a failedtarin a pipeline still lets the script continue and print "backup complete". - No overlap protection. A slow upload runs into the next schedule, and two jobs fight over the same archive or the same log.
-
Misread sync semantics.
aws s3 synccompares size and modification time. It does not delete destination objects unless you pass--delete. It also returns exit code 2 when some files were skipped (for example, unreadable files), which a naive script treats as success or ignores.
Reproduce the cron environment before debugging anything else:
env -i HOME="$HOME" PATH=/usr/bin:/bin /bin/sh -c '/opt/backup/backup.sh'
Run it as the cron user. If it fails there, you have found the class of bug.
Fix #1: Give the bash backup script an explicit environment
Stop depending on anything cron does not provide:
- Set
PATHat the top of the script and use absolute paths. - Prefer an instance profile or IAM role over static keys.
- If you must use a named profile, set
AWS_PROFILE, and setAWS_CONFIG_FILEandAWS_SHARED_CREDENTIALS_FILEif the files are not in the cron user's home directory. - Do not rely on IAM Identity Center (SSO) sessions for unattended jobs. They expire and cannot be refreshed without a browser login.
See the AWS CLI credential options in the official authentication documentation.
Crontab entry with the environment declared and output captured:
# /etc/cron.d/backup (system crontabs need a user field)
SHELL=/bin/bash
PATH=/usr/local/bin:/usr/bin:/bin
MAILTO=ops@example.com
# 02:15 daily; note the escaped percent signs if you use date formats here
15 2 * * * backup /usr/bin/flock -n /var/lock/backup.lock /opt/backup/backup.sh >> /var/log/backup/backup.log 2>&1
Watch out for the percent sign. In a crontab, an unescaped % is treated as a newline, so $(date +%F) directly in the cron line truncates the command. Put date logic inside the script, or write \%.
Also avoid cd-less relative paths. A script that does tar czf backup.tgz data/ resolves paths against cron's working directory. That is typically the cron user's home directory, but it is implementation-dependent, so cd explicitly or use absolute paths.
Fix #2: Fail loudly, lock, and clean up
Make every failure visible and stop overlapping runs. The flock -n wrapper in the cron line skips a run if the previous one still holds the lock. Inside the script, strict mode and a trap make failures non-silent and remove temporary archives.
A minimal skeleton that propagates errors and records the exit state:
#!/usr/bin/env bash
set -Eeuo pipefail # exit on error, unset vars, pipe failures
export PATH=/usr/local/bin:/usr/bin:/bin
SRC=/srv/app/data
BUCKET=s3://example-backups/app # replace with your bucket/prefix
STAMP=$(date -u +%Y%m%dT%H%M%SZ)
TMP=$(mktemp -d /var/tmp/backup.XXXXXX)
log() { printf '%s %s\n' "$(date -u +%FT%TZ)" "$*"; }
cleanup() { rm -rf "$TMP"; }
trap cleanup EXIT # always remove the temp dir
trap 'log "FAILED at line $LINENO"' ERR
log "start $STAMP"
tar -C "$SRC" -czf "$TMP/app-$STAMP.tgz" .
aws s3 cp "$TMP/app-$STAMP.tgz" "$BUCKET/archives/app-$STAMP.tgz" --only-show-errors
log "done"
Watch out for set -e inside conditionals and || or && chains. It is deliberately disabled there, so a failure inside an if test is not fatal. That also applies inside functions called from those contexts. Check critical commands explicitly when they sit there.
Uploading a single archive with aws s3 cp gives you one clear exit status. That is easier to alert on than a directory sync when the goal is a point-in-time backup.
Fix #3: Correct S3 sync behavior and rotate the logs
If you do mirror a directory, decide deliberately what deletions mean:
- Without
--delete,aws s3 synckeeps removed files in the bucket forever. - With
--delete, a mistakenly empty or unmounted source can wipe the destination prefix.
Guard against that and handle exit code 2:
# Refuse to sync from an unmounted or empty source
mountpoint -q /srv/app || { log "source not mounted, aborting"; exit 1; } # if SRC lives on its own mount
[ -n "$(ls -A "$SRC")" ] || { log "source empty, aborting"; exit 1; }
rc=0
aws s3 sync "$SRC" "$BUCKET/mirror/" --delete --only-show-errors || rc=$?
# rc 1 = transfer failed, rc 2 = some files skipped (e.g. permission denied)
[ "$rc" -eq 0 ] || { log "sync exit code $rc"; exit "$rc"; }
The || rc=$? form captures the status without tripping set -e or the ERR trap.
For history, enable S3 versioning and a lifecycle rule on the bucket rather than relying on script-side retention. Versioning alone does not stop someone holding credentials that can delete object versions. For ransomware protection, add S3 Object Lock as well. Flag details can change between CLI releases, so verify with the aws s3 sync reference.
Then stop the log from growing. A short job that appends with >> works fine with a standard logrotate rule:
# /etc/logrotate.d/backup
/var/log/backup/backup.log {
su backup backup
daily
rotate 14
compress
delaycompress
missingok
notifempty
create 0640 backup backup
}
The su line is needed because logrotate refuses to rotate in a directory writable by a non-root user.
If the writer is a long-running process that keeps the file descriptor open, add copytruncate. Otherwise it keeps writing to the renamed file.
Prevention
Backups fail silently by default, so build the checks into the process rather than trusting memory.
-
Test under cron conditions. Run the script with
env -iafter every change, and once as the actual cron user. - Alert on absence, not just errors. At the end of a successful run, send a ping to a dead-man's-switch service or write a success timestamp metric. A missing ping is the alarm, which also covers the case where cron never started.
-
Restore on a schedule. At least monthly, download the latest object and verify it with
tar -tzfor a checksum. A backup never restored is an assumption. -
Use least-privilege IAM. Scope the role to what each job needs:
- An archive-only job needs
s3:PutObjecton its prefix, pluss3:ListBucketon the bucket restricted with ans3:prefixcondition. - A mirror job using
--deletealso needss3:DeleteObject, scoped to the mirror prefix only. - With versioning on, deny
s3:DeleteObjectVersionso a buggy run only creates delete markers.
- An archive-only job needs
-
Consider systemd timers. On systemd hosts, a timer with
Persistent=trueruns missed jobs after downtime. Its logs land in the journal with exit status. - Keep scripts in version control and lint them with ShellCheck in CI.
For more Bash and infrastructure runbooks, browse the rest of kuryzhev.cloud.
Top comments (0)