Originally published on the Professional IT Services blog.
Every introduction to Kubernetes workload controllers gives the same three-line answer: Deployments for stateless apps, DaemonSets for one pod per node, StatefulSets for databases. That answer is correct, and it is also the part that never causes an outage. What causes outages is how each controller behaves when a node is drained, a volume refuses to attach, or a rollout hits a pod that never becomes healthy.
This deep dive answers the question from a production cluster rather than from the documentation. The worked example is our own K3s cluster on Hetzner Cloud: 5 nodes running 50 Deployments, 6 DaemonSets and 12 StatefulSets, with 26 persistent volumes. And the story that ties it together is the three-hour outage of 17 June 2026, when a routine node upgrade took every WordPress site on the cluster offline, and the workload controllers decided both how bad it got and how it was repaired.
One boundary first: this article is about which controller, and how it behaves when things break. Sizing — requests, limits, what the Galera pods and the DaemonSets actually consume — is covered separately in how we cut Kubernetes resource overhead by 50% using only built-in tools.
The 30-second decision tree
Ask three questions, in this order, and stop at the first yes.
- Does the pod work with the node itself — its logs, its metrics, its disks, its network, or traffic that arrives at that node? → DaemonSet.
- Does each replica need its own identity and its own data that must follow it when it is rescheduled — a database member, a message broker, a quorum peer? → StatefulSet.
- Everything else → Deployment. That includes an app that mounts a single persistent volume — but then choose the update strategy on purpose (see the next section).
In one sentence: a Deployment is for interchangeable pods, a StatefulSet is for pods that are not interchangeable, and a DaemonSet is for pods that belong to a node rather than to the application.
Deployments: the default, and the persistent-volume trap
A Deployment keeps a given number of identical pods running through a ReplicaSet. Pods get random names, any one can replace any other, and updates roll out gradually, with rollback built in. That is exactly right for web servers, APIs, workers and most of what a cluster runs, which is why 50 of our 68 workload controllers are Deployments.
Can a Deployment use persistent storage?
Yes. The previous version of this article said Deployments have ephemeral storage, and that was wrong: a Deployment can mount a PersistentVolumeClaim like any other pod. The trap is not the storage. It is the default update strategy combined with a ReadWriteOnce volume.
A Deployment's default strategy is RollingUpdate with maxSurge: 25% and maxUnavailable: 25%. Kubernetes rounds the surge up and the unavailability down. With replicas: 1, that means a surge of 1 and an unavailability of 0: the new pod starts before the old one stops. If the scheduler places the new pod on a different node, the ReadWriteOnce volume is still attached to the old node, the new pod sits in ContainerCreating with a Multi-Attach error, and the rollout stalls.
On our cluster, 15 Deployments mount a PVC. All 15 are ReadWriteOnce Hetzner volumes, and all 15 run a single replica. They include the WordPress sites, a mail stack whose five components share one 100 GiB volume, Nextcloud, Grafana and a newsletter server. Every one of them is a candidate for the trap.
After the June outage we reviewed all fifteen. The rule the review settled on: Recreate where it is necessary, RollingUpdate where it is feasible.
-
Recreatestops the old pod before starting the new one. Each rollout costs a few seconds of downtime, but it cannot deadlock on a volume. Three of the fifteen use it. -
RollingUpdateon one node. ReadWriteOnce means one node at a time, not one pod — the per-pod mode isReadWriteOncePod. Two pods on the same node can both mount the volume. So the WordPress sites keep their zero-downtime rollouts, and are kept on one node instead.
Our web chart does it with a pod-group label and a pod-affinity rule, so the new pod is drawn to the node where its predecessor, and any other pod sharing the volume, already runs:
strategy:
type: RollingUpdate
template:
metadata:
labels:
pod-group: {{ .Values.podAffinity.group }}
spec:
affinity:
podAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchExpressions:
- key: pod-group
operator: In
values:
- {{ .Values.podAffinity.group }}
topologyKey: kubernetes.io/hostname
The mail stack and Nextcloud go further and are pinned to a node with a nodeSelector.
If you copy this pattern, know its limit. preferred affinity is a hint, not a guarantee: when the node is under pressure, the scheduler is allowed to place the new pod elsewhere, and then you are back to the Multi-Attach error. required affinity or a nodeSelector gives the guarantee, at the price of a pod that stays Pending if that one node is full. If you have not thought about this trade-off for a particular app, Recreate is the safe default.
Does a DaemonSet really run on every node?
No. A DaemonSet runs one pod on every node it is allowed onto, and the difference is where the interesting decisions are.
Our cluster has a dedicated database node, tainted dedicated=mariadb:NoSchedule so that only the database lands there. This is what the six DaemonSets do with it:
| DaemonSet | Pods | Runs on the DB node? | Why |
|---|---|---|---|
fluent-bit (logs) |
5/5 | Yes, explicit toleration | No logs from the DB node is a blind spot |
loki-canary |
5/5 | Yes, explicit toleration | Same reason |
prometheus-node-exporter |
5/5 | Yes, tolerates every NoSchedule taint |
No metrics from the DB node is a blind spot |
hcloud-csi-node |
5/5 | Yes, tolerates every taint | Without the CSI node plugin, no volume attaches on that node |
ingress-nginx-controller |
4/5 | No, by design | The DB node serves no web traffic |
svclb-ingress-nginx-controller |
4/5 | No | Follows the ingress controller |
Same cluster, same taint, and the opposite answer is right for different DaemonSets. The observability agents must reach the tainted node — the database node is exactly where you want logs and metrics when something goes wrong. The ingress controller must not.
The detail that catches people: the DaemonSet controller automatically adds tolerations for the built-in node-condition taints (not-ready, unreachable, memory-pressure, disk-pressure, pid-pressure, unschedulable, and network-unavailable for host-network pods). It does not add tolerations for your own taints. Add a custom taint to a node, and every DaemonSet without a matching toleration silently stops covering it. There is no error, only a DESIRED count one lower than your node count. After tainting a node, kubectl get ds -A is the check.
Why run the ingress controller as a DaemonSet?
Because of how traffic reaches the cluster. A Hetzner load balancer forwards incoming traffic to all worker nodes, so every worker node needs an ingress controller to answer it. The chart supports both modes, and ours is set to DaemonSet:
controller:
# -- Use a `DaemonSet` or `Deployment`
kind: DaemonSet
The old version of this article also claimed that DaemonSets have "no built-in rolling update strategy (until Kubernetes 1.7+)". That is long obsolete: RollingUpdate is the default in the apps/v1 API, and all six of our DaemonSets use it with maxUnavailable: 1.
StatefulSets as they actually run in production
The textbook StatefulSet gives each pod a stable ordinal name (db-0, db-1, db-2), a stable DNS entry through a headless Service, and its own PersistentVolumeClaim from a volumeClaimTemplate, which follows the pod when it is rescheduled. By default it starts pods one at a time in ordinal order (podManagementPolicy: OrderedReady) and rolls updates in reverse order, with a partition to stage them.
Our production database is a three-node MariaDB Galera cluster, and it shows how differently a real one is configured. It is not a hand-written StatefulSet. It is declared as a MariaDB custom resource, and mariadb-operator generates the StatefulSet from it:
apiVersion: k8s.mariadb.com/v1alpha1
kind: MariaDB
spec:
replicas: 3
galera:
enabled: true
affinity:
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchExpressions:
- key: app.kubernetes.io/instance
operator: In
values:
- mariadb-galera-new
topologyKey: kubernetes.io/hostname
podDisruptionBudget:
minAvailable: 2
The ordering people choose StatefulSets for is switched off
The generated StatefulSet runs with podManagementPolicy: Parallel and updateStrategy: OnDelete. The ordered, one-at-a-time behaviour that most articles list as the main reason to use a StatefulSet is turned off — and the operator does the ordering itself, rolling the replicas first and the primary last.
The reason is how Galera works. A Galera member that cannot see its peers cannot form a quorum, so "start pod-0, wait until it is Ready, then start pod-1" can stall a full-cluster restart. The operator starts the members together, and applies the order that actually matters for a database — never restart the node taking writes until the others are updated — in its own logic.
The rest of the configuration is where the operational safety lives:
-
Required anti-affinity on the hostname, so no two members share a node, plus a
nodeSelectorand a toleration that admit them to the database nodes only. -
A PodDisruptionBudget of
minAvailable: 2. A three-node Galera needs two nodes for quorum, so a node drain may take down at most one member at a time. -
persistentVolumeClaimRetentionPolicy: Retainfor both deletion and scale-down. Deleting the StatefulSet or scaling it down never deletes the data. In June this was what made the repair possible. -
Two
volumeClaimTemplatesper pod: a 20 GiB data volume and a 100 MiB Galera state volume. The 100 MiB claims are bound at 10 GiB, because Hetzner's minimum volume size is 10 GB — three times over, for about 300 MB of state. - Volumes are tied to a location. A Hetzner volume attaches only to a server in the same location. A PVC created in one datacenter cannot follow its pod to another one; moving it means copying the data into a new PVC.
17 June 2026: what the controllers did during a three-hour outage
The Galera cluster in production at the time was the previous one: a Bitnami chart StatefulSet with the default ordered pod management. This is how a routine maintenance task became a total outage.
- A routine worker-node upgrade drained a node.
- Galera lost its Primary component. All three members went non-Primary, and every WordPress site on the cluster returned "Error establishing a database connection". Contributing factors, as the runbook records them: one member was running in a different datacenter, the pods sat at about 99% of a 5 GiB memory limit, and heavy database dumps were running at the same time.
- The CSI pods could not be scheduled either. No CSI, no volume attachments — and from that point everything that needed a persistent volume was down, not just the database.
-
Recovery started by bootstrapping node-0 with an empty cluster address (
gcomm://), so it formed a new Primary component on its own. -
The StatefulSet's rolling update was stuck. Under
OrderedReady, the controller will not roll past a pod that never becomes Ready; the Kubernetes documentation says that in this state you have to delete the broken pods by hand. The fix was to delete the StatefulSet without deleting its pods or PVCs and recreate it:
kubectl delete statefulset <name> --cascade=orphan
-
Three-node HA came back through a staged rollout with
partition=1. Node-0 kept serving untouched, nodes 1 and 2 rejoined through a full state transfer (SST), and node-0 rolled last. The memory limit went from 5 GiB to 7 GiB.
Total downtime: about three hours. No data was lost.
What changed afterwards
The Bitnami image in use had been moved to the frozen bitnamilegacy repository, with no further security updates, so the cluster was moved to mariadb-operator. The first cutover attempt — a full dump and reload under a write freeze — took about 60 minutes on a three-node Galera, which is far too long to freeze writes. It was redone the next day as asynchronous GTID replication from the old cluster to the new one, followed by a write freeze of seconds and a switch of the Service selector.
The database's DNS name never changed. The old mariadb-galera Service now selects the new primary, so not a single tenant configuration file had to be touched. At the same time, the database and several application volumes were moved from the old datacenter to the main one — by copying the data into new PVCs, since a volume cannot move in place.
Lessons that apply to any cluster
- A node drain is decided by your PodDisruptionBudgets, taints and tolerations — check what each controller will do before upgrading nodes, including the CSI components.
- Know what ordered pod management does with a pod that never becomes Ready. It waits. The rollout will not repair itself.
-
--cascade=orphanplus aRetainPVC policy lets you rebuild a controller without touching the running pods or the data. - Keep the Service name stable. The DNS name is the contract with every client of the database; everything behind it can be replaced.
-
Move data safely: never dump inside the database pod (it will run out of memory), never stream large data through the kube-apiserver (a small control-plane node will fall over), copy pod to pod inside the cluster with a bandwidth cap (
pv -L 8m), one database at a time, withset -o pipefail.
Key differences at a glance
| Deployment | DaemonSet | StatefulSet | |
|---|---|---|---|
| Manages | Interchangeable replicas | One pod per eligible node | Replicas with stable identity |
| Pod names | Random hash | Random hash, one per node | Ordinal: name-0, name-1
|
| Placement | Scheduler decides | Every node it tolerates | Scheduler decides; anti-affinity usually required |
| Storage | Any PVC; shared RWO volumes need care | Usually node-local paths | Own PVC per pod via volumeClaimTemplates
|
| Default update |
RollingUpdate, 25% surge / 25% unavailable |
RollingUpdate, maxUnavailable: 1
|
RollingUpdate in reverse ordinal order, with partition
|
| Ordering | None | None |
OrderedReady by default; clustered DB operators often use Parallel
|
| Typical failure | Multi-Attach error on a single-replica RWO rollout | Silently skips nodes with custom taints | Rollout stuck on a pod that never becomes Ready |
| Use for | Web apps, APIs, workers | Logs, metrics, CSI, ingress on every node | Databases, brokers, quorum-based systems |
Final thoughts
Choosing the controller takes thirty seconds with the decision tree above. The work is in the behaviour around it: the update strategy of a Deployment with a volume, the tolerations of every DaemonSet, and the pod management, disruption budget and retention policy of every StatefulSet. On our cluster, each of those was either a cause of the June outage or part of its repair.
Keeping the resulting cluster efficient is the other half: our production K3s audit — how we cut Kubernetes resource overhead by 50% using only built-in tools — shows what over-requested StatefulSets and cluster-wide DaemonSets actually cost once they are measured.
Top comments (1)
Some comments have been hidden by the post's author - find out more