The cron job ran at 3am. By 3:02 the server was unreachable.
Not slow. Gone. SSH timing out, no response, nothing in the logs because the thing that writes the logs had itself been killed.
The cause was not a bug. It was arithmetic. A batch job that looked harmless on my laptop quietly needed more memory than the box had, and the kernel did what kernels do: it picked a victim and killed it.
Why cheap servers fail differently
On a big machine, memory pressure is a performance problem. On a 2GB box, it is an availability problem. You do not get "slow". You get "dead", and you get it at the worst possible time, which is always 3am.
The specific killer is almost always concurrency. Running several image or file operations at once multiplies peak memory, and the peak is what matters, not the average. A job that uses 300MB on average and 2.4GB at peak will take down a 2GB server every single time.
The fix, in order of how much it helped
One process at a time. This alone ended the crashes. It felt slower and it was not, because the crashed runs were costing me whole nights.
Measure the peak, not the average. I started printing the worst-case memory for each step. The average is a comfortable lie; the peak is the truth.
Free memory between steps. In Python, large objects do not always go back to the OS the moment you drop the reference. Asking the garbage collector to run between heavy steps keeps the peak down.
Cap the input size. The job that crashed was processing a batch bigger than it needed to. Chunking the work into smaller pieces lowered the peak without changing the total.
The rules I now follow on any small box
- One heavy job at a time. No exceptions.
- Know the peak memory of every step before you schedule it.
- Chunk the input so a single run can never blow the budget.
- Make the job idempotent, so a run that dies can be safely repeated.
- Watch swap. Heavy swap use is the early warning before the crash.
The uncomfortable lesson
I spent a week blaming the code. The code was fine. I was asking a 2GB server to behave like a 16GB one, and no amount of cleverness fixes that.
The boring answer, one process and a smaller batch, worked on the first try. Cheap infrastructure is not a limitation you outsmart. It is a budget you respect.
If you run pipelines like this, the free sampler here shows what mine produces (transparent PNG, commercial-use, no email wall): Free cute sticker sampler
The scripts, memory guards and all, are collected in The Pipeline Starter Kit.
Top comments (1)
The subtle trap with Python on small VPS boxes is that gc.collect() cleans up Python object references, but glibc malloc arenas frequently hold onto the pages and never release them back to the Linux kernel. If fragmentation keeps a high watermark, the next step starts from an already inflated baseline.
Putting the memory-heavy batch step inside a dedicated worker subprocess that exits when finished guarantees the kernel immediately reclaims every byte of RSS. It also avoids swap thrashing, which is usually what hangs the SSH daemon for ten minutes before the OOM killer even steps in.