On September 6 a disk alert fired on the server that runs my voice bots: 75% full. The culprit was /var/log/auth.log, 763 MB after five days, on a box where it normally grows about 0.7 MB a week. Nearly every line started with audispd:.
I knew why. Six days earlier I had switched on auditd's syslog plugin across the fleet so a central collector could receive the audit trail. I compressed the file, wrote the cause down, and moved on.
This morning I measured it on all eight servers. The growth turned out to be the smaller finding. The archive I built to keep audit logs for 24 months had been saving about a third of each day on my busiest server, from the week I set it up.
The setup
Eight Ubuntu 24.04 servers running auditd 3.1.2, with 22 to 25 rules each. Most rules are file watches, -w <path> -p wa, on config files and on each app's directory. Israel's data security regulations require keeping access logs for 24 months. auditd knows nothing about months. Out of the box it keeps a ring:
# /etc/audit/auditd.conf
max_log_file = 8 # MB
num_logs = 5
max_log_file_action = ROTATE
Five files of 8 MB each. When the ring is full, the oldest file is deleted. So on August 25 I added a cron job at 02:17 that forces a rotation, gzips every rotated file into /var/log/audit-archive/, and deletes archives older than 730 days. Each archive's filename carries a hash of its content, so a file that moves from audit.log.1 to audit.log.2 isn't archived twice.
On August 31 I enabled the syslog plugin:
# /etc/audit/plugins.d/syslog.conf
active = yes
args = LOG_INFO LOG_AUTHPRIV
LOG_AUTHPRIV is why the copies land in auth.log. Ubuntu's rsyslog routes auth,authpriv.* there, and logrotate keeps that file for 104 weeks.
What auth.log looks like now
The same seven days of the week, before and after August 31, on six of the eight servers. The other two are dedicated to a client, so their numbers stay out of this post.
| Server | Aug 24–30 | Sep 21–27 | Growth | audispd share |
|---|---|---|---|---|
| database | 2.3 MB | 734 MB | ×324 | 99.7% |
| website | 1.4 MB | 729 MB | ×524 | 99.6% |
| voice | 0.6 MB | 275 MB | ×441 | 99.1% |
| WhatsApp gateway | 1.0 MB | 96 MB | ×101 | 98.0% |
| n8n | 1.3 MB | 88 MB | ×70 | 95.5% |
| chatwoot | 1.4 MB | 1.5 MB | ×1 | 0% |
The last row is a control group: the plugin has been off on that server since September 8.
Who is writing all of this?
Not people. aureport shows which rule and which executable produced the events in the local ring:
sudo aureport --input-logs -k --summary # events per rule key
sudo aureport --input-logs -x --summary # events per executable
Keep --input-logs. When stdin isn't a terminal, aureport reads stdin instead of the log files and prints <no events of interest were found>. That cost me ten confused minutes this morning.
On the database server, 35,247 of the 35,659 events that matched a rule came from a single watch, -w /opt/supabase -p wa, and the executable behind 35,082 of them was postgres. The app directory contains the database's data volume, so every file Postgres opens for writing (paths like base/5/2601) becomes an audit event.
On the website, 19,120 events hit the site's watch, and rsync (14,988) and rm (4,064) account for nearly all of them. That's the deploy. One deploy last night, between 23:07 and 23:12, produced 13,151 events made of 70,516 records.
On the voice server it's the same picture: 24,687 events on the app's watch, almost all from rsync. The deploy there rewrites about 30,000 files on every run, changed or not.
One write to a watched path isn't one line. It's a SYSCALL record plus CWD, one or two PATH records and PROCTITLE, 5.4 records per event on average in that deploy. The syslog plugin also forwards the EOE (end of event) record, which auditd doesn't keep in its own log, so auth.log gets one more line per event than audit.log does.
The ring runs out before the archive runs
A 40 MB ring is fine as long as 40 MB lasts longer than the gap between archive runs. Here's how long it lasts, from aureport --input-logs -t:
- database: the five files cover 13.5 hours
- website: around last night's deploy, three 8 MB files filled in 3 min 48 s, then 58 s, then 17 s
- chatwoot: the five files cover almost three days
The archive ran once a night. Anything that left the ring before 02:17 was deleted before anyone saved it.
So I measured the archive itself: for every archived file, the first record's timestamp and the file's mtime, merged into intervals and compared with each calendar day. Twenty-nine days, August 30 to September 27:
| Server | Mean coverage | Worst day | Complete days |
|---|---|---|---|
| database | 33.6% | 15.0% | 0 of 29 |
| website | 52.6% | 9.4% | 9 of 29 |
| voice | 75.0% | 22.0% | 16 of 29 |
| chatwoot | 94.0% | 16.4% | 26 of 29 |
On the database server, the "24-month archive" holds 234 of the last 696 hours. On the website, deploy days drop to about 10%, because a deploy late in the evening pushes everything before it out of the ring. chatwoot's three incomplete days were maintenance days: a hardening pass, a Docker update and a fleet-wide reboot.
The plugin didn't cause this. On August 30, the day before I enabled it, the database server's archive already held 28.4% of the day. Postgres had been filling the ring since the archive's first week.
The noise was the backup
The audispd lines in auth.log, the ones I spent September 6 compressing, are kept for 104 weeks. Since August 31 they have been the most complete local copy of the audit trail on these servers.
They're not a perfect copy. I counted SYSCALL records, one per event, by the timestamp inside each record, in both logs, for 23:04 to 02:00 last night:
-
audit.log: 19,177 events -
auth.log: 18,912 events (98.6%)
The missing 1.4% left a trail too. The plugin writes to /dev/log, which on Ubuntu 24.04 is journald's socket, and journald forwards to rsyslog. During that window journald logged four Forwarding to syslog missed N messages lines, 2,789 messages in total. The journal's own copy fared worse: between 23:00 and 23:20 its rate limiter logged Suppressed … messages from auditd.service six times, 101,004 messages in all.
What I changed today
The smallest change that stops the loss is to archive more often without forcing a rotation every time. The script got a flag and a lock:
exec 9>/run/audit-archive.lock
flock -w 120 9
if [ "${1:-}" != "--no-rotate" ]; then
/sbin/auditctl --signal rotate >/dev/null 2>&1 || pkill -USR1 auditd 2>/dev/null || true
sleep 2
fi
The cron file got a second line:
17 2 * * * root /usr/local/bin/audit-archive.sh >/dev/null 2>&1
*/15 * * * * root /usr/local/bin/audit-archive.sh --no-rotate >/dev/null 2>&1
The nightly run still forces a rotation, so quiet servers get a daily file. The quarter-hourly run only picks up files auditd already rotated by size, which is exactly the part that used to fall off the end. The lock keeps the 02:15 and 02:17 runs from writing the same temporary file. I rolled it out to all eight servers with the old files backed up. The first run on the database server archived three files that would have been gone before tonight's 02:17.
It doesn't close every hole. A burst bigger than about 32 MB inside fifteen minutes, like two website deploys back to back, can still push a file out before the archive sees it. The durable fixes all mean writing less, and I haven't made them yet:
-
Stop watching the database's data directory. On August 25 I narrowed a
/var/lib/dockerwatch to/var/lib/docker/containersfor the same reason: Postgres block writes say nothing about who accessed what. The/opt/supabasewatch needs the same change. - Make the deploys skip unchanged files. Thirty thousand rewritten files per deploy is an rsync problem, and auditd is only reporting it faithfully.
-
Give the syslog plugin its own facility,
LOG_LOCAL6with its own file, so auth.log goes back to being about logins.
Check your own servers
Three commands:
# how big is the ring?
grep -E '^(max_log_file|num_logs|max_log_file_action) ' /etc/audit/auditd.conf
# how many hours does it hold right now?
sudo aureport --input-logs -t
# who is filling it?
sudo aureport --input-logs -x --summary | head
If the second command shows fewer hours than the gap between your archive runs, your retention policy exists only in a document. I build and run automation infrastructure for small businesses at Achiya Automation, including the voice server in this post. This is the second time in five weeks that a retention number I had written down turned out to be one I had never measured.
A question for anyone who archives auditd logs for compliance: does your archive job run on rotation or on the clock, and how many hours does the ring actually hold on your busiest server? I'd especially like to hear from anyone running max_log_file_action = keep_logs with a separate mover script, and whether that held up the first time the disk got tight.
Top comments (2)
If you'd rather measure your own archive than trust the retention policy, this is roughly the check I ran on each server (as root). It prints how much of each UTC day the archived files actually cover:
Adjust the glob to wherever your archive lives. It treats each file as covering its first record's timestamp through its mtime, which holds as long as nothing touches the files after they're written. Any day under 100% is a day your "24 months" doesn't actually include.
hold up. green uptime tiles are not a usage receipt.
1 cut: when the cloud invoice fight opens, can a buyer GET a signed meter tip of what ran, or only another vendor seal?
receipts > seals. #marker1328-obs