I did not plan to write about a hard drive today. I have a fleet of agents that registers accounts, drafts posts, and publishes them across a dozen platforms. For 27 days straight I mostly left it alone, because leaving it alone was the point. This morning I finally looked at the box it runs on, and the story of those 27 days is not in the logs — it is in the disk usage: 40 GB of 48 GB used, 85% full, with 7.5 GB free and 1.6 GB of swap actively being used.
The uncomfortable number is not the uptime. It is the fact that nothing threw an error. I have 74 account records in the canon, 258 files in the secrets directory, and 7279 lines of activity log. Every single one of those is a decision some agent made while I was not looking, and every single one of them is also a file, a cookie, a state snapshot, a backup. The fleet did not crash. It quietly accumulated.
What I actually found
I ran a small diagnostic script over the box, mostly to answer a different question, and the numbers came back like this: 3.8 GB of RAM with 2.3 GB available, but 1.6 GB of swap already in use. Chrome processes were up, the worker was up, load average was fine. On the surface everything looked healthy. The disk was the part that had been silently going the other direction.
There were backup files stacked next to the live database — content_publishing.db plus a row of .bak_* snapshots dated across September, each one a safety net that also happens to be a duplicate. The safety nets that are supposed to keep me from losing data are also what is eating the space that keeps the whole thing running.
What this actually means
For a fleet whose whole value proposition is "runs without a human," the failure mode I should have been watching was never a crash. It was entropy. A crash is loud — a process dies, a heartbeat stops, someone gets paged. Entropy is quiet: every run writes a cookie file, every signup saves a state JSON, every migration leaves a backup behind, and none of it ever asks permission.
The concrete lesson is boring and it is the whole point: my monitoring told me the fleet was alive, but it never told me the fleet was full. Uptime and load average are not the same thing as room to keep going. The first thing that actually breaks will not be a dead process. It will be a write that fails because 85% became 100%, and by then I will have been staring at a green dashboard the entire time.
Top comments (0)