DEV Community

Charles Hartmann
Charles Hartmann

Posted on AI-assisted

Backup, verify, destroy: how I deleted 50 Proxmox guests in one night and could undo any of them

If your Proxmox guest list is mostly grey "stopped" icons, this is the routine I used this week to delete 50 of them and get about 700 GB back, with every single one recoverable.

First, find out what's actually unused

Stopped is not the same as unused. A template, or a plain container someone turned into a template, is always stopped and may be the base of linked clones. Deleting it takes every clone with it.

For every candidate ID, grep the cluster's configs for a disk that references it as a base:

grep -l "base-<id>-" /etc/pve/nodes/*/*/*.conf
Enter fullscreen mode Exit fullscreen mode

If anything comes back, that guest is a base image. I found two bases that six running containers depended on. Tag those keep and move on.

Then, for each real candidate: backup, verify, destroy

Back it up first. --mode stop gives a consistent image of a guest that is already stopped anyway:

vzdump <id> --storage <backupstore> --compress zstd --mode stop
Enter fullscreen mode Exit fullscreen mode

Verify the archive before you trust it. A backup you haven't at least integrity-checked is a rumor:

zstd -t /path/to/vzdump-lxc-<id>-*.tar.zst
# VMs:
zstd -t /path/to/vzdump-qemu-<id>-*.vma.zst
Enter fullscreen mode Exit fullscreen mode

Only then destroy it. --purge matters: without it the old backup job keeps referencing the dead ID and fails on its next run.

pct destroy <id> --purge
# VMs:
qm destroy <id> --purge
Enter fullscreen mode Exit fullscreen mode

Watch the backup storage's free space as you go. I went in order of "would be saddest to lose" and stopped when the share got tight.

Afterwards: the leftovers

  • Orphan disks: pvesm list <storage> entries that no config references.
  • Old ISOs and CT templates sitting in local.
  • Snapshots older than a year.

That was another ~100 GB.

Two gotchas

Running the vzdumps at full speed made the host's busy containers laggy, so I added --bwlimit to the vzdump command.

On a two-node cluster, the second node kept losing quorum (it's on Wi-Fi), so every destroy on it had to be retried while quorate. A QDevice fixed that properly.

The checklist

  1. List every stopped guest.
  2. Grep the configs for base-<id>-; tag any hit keep.
  3. For the rest: vzdump with zstd, zstd -t the file, then destroy with --purge.
  4. Check the backup job still covers everything that's left (mine was set to "selected VMs" years ago and covered half the guests).
  5. Clean up orphan disks, old ISOs, old snapshots.

This is Chapter 4 of Homelab the Right Way (Proxmox, OPNsense & Home Assistant Without the Pain). The chapter and the weekend checklist are free: https://payhip.com/b/qon0W. The full 40-page book is $19: https://payhip.com/HomelabtheRightWay.

Top comments (0)