My etcd database was 508 MB on disk and only 84 MB of it held live data. The other 400-odd megabytes were free pages that etcd had no intention of handing back to the filesystem. Defragmenting brought it back down to 84 MB. That part took seconds. Doing it without the control plane noticing took more care.
If you run your own control plane (kubeadm on bare metal, Proxmox VMs, anything where nobody else looks after etcd for you), this is your job. Managed Kubernetes providers do it in the background. On a homelab or on-prem cluster, nobody does it unless you set it up, and nothing tells you it's needed until you hit a quota alarm or start seeing odd latency.
Why etcd grows and never shrinks
etcd stores data in bbolt, a B+tree in a single memory-mapped file. Every write to a Kubernetes object creates a new revision. Lease renewals, Endpoints updates, Event objects and controller status patches all add revisions. The kube-apiserver compacts old revisions every 5 minutes by default (--etcd-compaction-interval), so history doesn't pile up forever.
The catch is that compaction only marks pages as free inside the bbolt file. The file stays the same size. bbolt reuses those free pages for new writes, so the file doesn't grow without limit, but after a burst of churn (a big rollout, a namespace deletion, a controller stuck in a status-update loop) the file keeps its high-water mark.
That causes three problems:
-
The quota counts file size, not live data.
--quota-backend-bytesdefaults to 2 GiB. When the file hits that limit, etcd raises aNOSPACEalarm and goes read-only for writes. The apiserver starts rejecting creates and updates, even if 80% of the file is empty pages. -
Snapshots copy the whole file. A
snapshot saveof a fragmented database writes a fragmented snapshot, so bloat gets multiplied by however many backups you keep. - Bigger mmap, slower startup. A member that has to restart and load a bloated database takes longer to rejoin.
Defragmentation rewrites the bbolt file with only the live pages and releases the rest back to the filesystem. It's the only way to shrink the file short of restoring from a snapshot.
Know when it's actually needed
Don't defrag on a timer just because you can. Defrag when the gap between on-disk size and in-use size is large. etcdctl endpoint status reports both in its JSON output:
export ETCDCTL_API=3
export ETCDCTL_CACERT=/etc/kubernetes/pki/etcd/ca.crt
export ETCDCTL_CERT=/etc/kubernetes/pki/etcd/server.crt
export ETCDCTL_KEY=/etc/kubernetes/pki/etcd/server.key
ENDPOINTS="https://10.0.0.11:2379,https://10.0.0.12:2379,https://10.0.0.13:2379"
etcdctl --endpoints="$ENDPOINTS" endpoint status -w json | jq -r '.[] |
"\(.Endpoint) total=\(.Status.dbSize / 1048576 | floor)MiB " +
"inuse=\(.Status.dbSizeInUse / 1048576 | floor)MiB " +
"ratio=\((.Status.dbSizeInUse / .Status.dbSize * 100) | floor)% " +
"leader=\(.Status.leader == .Status.header.member_id)"'
Before my defrag, the output looked like this (endpoints anonymized):
https://10.0.0.11:2379 total=508MiB inuse=84MiB ratio=16% leader=true
https://10.0.0.12:2379 total=497MiB inuse=84MiB ratio=16% leader=false
https://10.0.0.13:2379 total=502MiB inuse=84MiB ratio=16% leader=false
The in-use size is almost identical across members because they all hold the same data. The on-disk size differs a bit because each member's bbolt file fragments on its own. A ratio of 16% means five-sixths of the file is dead space.
My rule of thumb: act when the in-use ratio drops below 50% and the total is large enough to matter (more than about 100 MiB). A 40 MiB database at 45% isn't worth a maintenance window. The research brief's version of this check used "DB size vs. total keys", which works as a rough signal, but dbSizeInUse measures the thing you care about directly.
The zero-disruption part: one member at a time, followers first
Defragmentation blocks the member it's running on. While bbolt rewrites the file, that member can't serve reads or writes. On a small database this takes a second or two. On a multi-gigabyte one it can take much longer. The etcd and Kubernetes docs both say the same thing: defrag one member at a time, and never touch more than one at once.
Two things follow from that:
-
Don't use
etcdctl defrag --cluster. It goes through every member in the member list back to back, with no pause and no health check between them. If one member is slow to recover, the next one is already blocked, and in a three-member cluster two members down means no quorum. - Do the leader last. A defrag on a follower barely matters: the leader keeps committing with the other follower, and the defragged follower catches up afterwards. A long defrag on the leader can stall it past the followers' election timeout (1000 ms by default). That triggers an election, and every in-flight apiserver request to etcd waits on it.
Here's the script I run from a control plane node. It finds the leader, defrags each follower, checks health after each one, and handles the leader last:
#!/usr/bin/env bash
set -euo pipefail
source /etc/etcd-maint.env # exports ETCDCTL_* certs and ENDPOINTS
leader=$(etcdctl --endpoints="$ENDPOINTS" endpoint status -w json \
| jq -r '.[] | select(.Status.leader == .Status.header.member_id) | .Endpoint')
[[ -n "$leader" ]] || { echo "no leader found, aborting"; exit 1; }
followers=$(tr ',' '\n' <<<"$ENDPOINTS" | grep -vxF "$leader")
for ep in $followers "$leader"; do
echo "== defrag $ep"
etcdctl --endpoints="$ep" --command-timeout=120s defrag
# Every member must be healthy before we touch the next one
etcdctl --endpoints="$ENDPOINTS" endpoint health
sleep 30
done
etcdctl --endpoints="$ENDPOINTS" alarm list
Some notes on the choices:
-
--command-timeout=120smatters. The default is 5 seconds, and on a large database the client gives up while the server is still defragmenting. You get an error, the defrag finishes anyway, and you're left unsure what state things are in. - The health check runs against all endpoints, not just the one you defragged. Under
set -e, one unhealthy member stops the script before the next defrag starts. That's the whole safety mechanism, so don't remove it. - The 30-second sleep lets the defragged member catch up on raft entries it missed. It isn't strictly required, but it's cheap.
-
alarm listat the end catches a leftoverNOSPACEalarm (more on that below).
On a kubeadm cluster, if you don't want to install etcdctl on the host, you can kubectl exec into each etcd-<node> static pod and run the same commands. The certificate paths are the same because kubeadm mounts /etc/kubernetes/pki/etcd into the pod.
Handling the leader: move it or accept the election
If the leader stalls, leadership can move somewhere else, and that's what happened in my maintenance window. After I defragged the leader, a different node held leadership and the raft term had gone up. Nothing broke. The apiserver retried and the cluster stayed healthy. It still made me want to check afterwards whether the cluster had settled or was flapping.
There's a cleaner option: move leadership off the node yourself, then defrag it as a follower.
# Pick a follower that's already been defragged as the new leader
new_leader_id=$(etcdctl --endpoints="https://10.0.0.12:2379" endpoint status -w json \
| jq -r '.[0].Status.header.member_id | tostring')
# move-leader has to be sent to the current leader's endpoint
etcdctl --endpoints="$leader" move-leader "$(printf '%x' "$new_leader_id")"
move-leader takes the member ID in hex, which is the format etcdctl member list prints. The JSON output gives it in decimal, hence the printf. There's also a jq trap here: member IDs are 64-bit, and older jq versions convert large integers to doubles and lose precision. If the hex ID looks wrong, take it straight from etcdctl member list -w table rather than trusting the conversion.
Once leadership has moved, the old leader is a follower, and your defrag loop treats it like any other follower. There's no surprise election.
Whichever route you take, confirm the cluster is stable afterwards:
# Run twice, a minute apart. Same leader and same raft term = settled.
etcdctl --endpoints="$ENDPOINTS" endpoint status -w json | jq -r '.[] |
"\(.Endpoint) leader=\(.Status.leader == .Status.header.member_id) term=\(.Status.raftTerm)"'
kubectl get --raw='/readyz?verbose' | grep etcd
If the raft term keeps rising between the two runs, leadership is still bouncing. Stop and investigate (disk latency is the usual cause) before you blame the defrag.
The NOSPACE alarm doesn't clear itself
If you're defragging because you already hit the quota, there's one more step. Defragmentation frees the space, but the NOSPACE alarm stays raised until you disarm it:
etcdctl --endpoints="$ENDPOINTS" alarm list
# memberID:1234567890 alarm:NOSPACE
etcdctl --endpoints="$ENDPOINTS" alarm disarm
etcdctl --endpoints="$ENDPOINTS" alarm list # should print nothing
Until you disarm it, etcd keeps refusing writes even though the file is now a fraction of its old size. This is the step people miss when they defrag during an outage and then wonder why kubectl apply still fails.
The side effect nobody mentions: snapshot sprawl
This is the part the research brief called the "total cost of maintenance", and it's the lesson I'd pass on first. You should take a snapshot before any defrag. Defrag has been reliable in my experience, but "reliable" isn't a backup strategy. On a control plane with a small boot disk (mine are in the 30 GB range), those snapshots add up quickly, especially snapshots of a fragmented database. You fix etcd bloat and get DiskPressure on the control plane node instead, and the kubelet starts evicting things.
The fix is to prune snapshots automatically, on the same schedule that creates them. A systemd timer works well:
# /etc/systemd/system/etcd-snapshot.service
[Unit]
Description=etcd snapshot with retention
[Service]
Type=oneshot
EnvironmentFile=/etc/etcd-maint.env
ExecStart=/bin/sh -c 'etcdctl --endpoints=https://127.0.0.1:2379 \
snapshot save /var/backups/etcd/etcd-$(date +%%Y%%m%%dT%%H%%M).db'
ExecStartPost=/usr/bin/find /var/backups/etcd -name "etcd-*.db" -mtime +3 -delete
# /etc/systemd/system/etcd-snapshot.timer
[Unit]
Description=Daily etcd snapshot
[Timer]
OnCalendar=*-*-* 03:15:00
RandomizedDelaySec=15m
Persistent=true
[Install]
WantedBy=timers.target
Enable it with systemctl enable --now etcd-snapshot.timer. Three days of retention is enough locally because the snapshots also go off-box. A snapshot that only lives on the control plane disk doesn't protect you if you lose that disk. In my setup cluster-level backups go to object storage (the Velero + MinIO setup covers that side), and the control plane VMs themselves are backed up at the hypervisor level.
A useful side effect: defragging before your snapshot window makes every snapshot after it smaller, because the snapshot size follows the bbolt file size. In my case that's 84 MB per snapshot instead of 508 MB, which changes how many days of retention a small disk can hold.
Alert on the ratio, not the size
I found my bloat after the fact, while chasing pod restarts that looked unrelated. Nothing had alerted. That's the observability gap: etcd fragmentation doesn't show up as a failing health check, just as a number drifting upward on a metric nobody watches.
etcd exposes both sizes as Prometheus metrics. The etcd-mixin (bundled with kube-prometheus) already ships a rule for this, and if you're writing your own, it looks like this:
groups:
- name: etcd-maintenance
rules:
- alert: EtcdDatabaseHighFragmentation
expr: |
(last_over_time(etcd_mvcc_db_total_size_in_use_in_bytes[5m])
/ last_over_time(etcd_mvcc_db_total_size_in_bytes[5m])) < 0.5
and etcd_mvcc_db_total_size_in_use_in_bytes > 104857600
for: 1h
labels:
severity: warning
annotations:
summary: "etcd {{ $labels.instance }} is >50% free pages; schedule a defrag"
The size guard (100 MiB in use) keeps the alert quiet on small databases, where a bad ratio doesn't matter. The one-hour for rides out short-lived churn during rollouts. Severity is warning, not critical: this is "schedule maintenance this week", not "wake someone up". Pair it with a separate rule on etcd_mvcc_db_total_size_in_bytes approaching etcd_server_quota_backend_bytes, and give that one a higher severity, because it means writes are about to stop. I wrote more about keeping rules like these actionable in Prometheus Alerting Rules That Don't Cry Wolf.
If your control plane is scraped over HTTPS with client certs (kubeadm's default puts metrics on 127.0.0.1:2381 over plain HTTP), check that Prometheus can reach it before you trust the absence of an alert. A missing metric doesn't fire an alert, which is how the gap stays open.
Gotchas and alternatives I considered
Offline defrag with etcdutl. If a member is stopped (for example during a node rebuild), etcdutl defrag --data-dir /var/lib/etcd defragments the files directly without a running server. It's fine for a member that's down anyway. It's the wrong tool for a live cluster, because stopping a member to defrag it offline costs more availability than an online defrag.
Defrag doesn't compact. If compaction isn't running (a custom apiserver flag, or a standalone etcd without --auto-compaction-retention), defrag reclaims almost nothing, because the old revisions are still live data. Check that dbSizeInUse is reasonable first. If in-use is huge, the problem is compaction or a runaway writer, not fragmentation.
Find the churn source. A database that re-fragments within days means something is writing constantly. Event objects with long TTLs, a controller hot-looping on status updates, and oversized ConfigMaps rewritten on every reconcile are the usual suspects. Defrag treats the symptom. etcd_debugging_mvcc_keys_total and apiserver request metrics grouped by resource will point you at the cause.
Disk latency makes everything worse. Defrag writes a whole new copy of the database. On slow or shared storage, that write burst can push fsync latency past what raft tolerates, and then you get elections you didn't ask for. If your control plane VMs share a disk with noisy neighbors, run defrags in quiet hours.
Scheduled automatic defrag. Some people run a nightly defrag on cron. I prefer letting the alert drive it: defrag is cheap but not free, and running it unconditionally hides the churn signal you'd otherwise notice. If you do automate it, use the follower-first script with health gates, never --cluster.
Wrap-up
Defragmentation is safe when you treat it as a rolling operation: snapshot first, one member at a time, followers before the leader (or move leadership off first), health-gate every step, disarm any alarm, and then check that leadership has settled. The command itself is the easy part. The work is in the process around it, especially snapshot retention on small control plane disks, which is how cleaning up one problem turns into the next one.
Use this when the in-use ratio drops below half on a database big enough to matter, when you've hit NOSPACE, or before you take a snapshot you plan to keep for a long time. Set up the fragmentation alert so you hear about it from a warning rather than from a failing kubectl apply. If you're running control planes like this and want a second opinion on the operational side, that's the kind of work I do through GuatuLabs. For more on how the control plane VMs fit into the wider cluster, see Building a Production Homelab.
Top comments (0)