<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Csaba Ajtony</title>
    <description>The latest articles on DEV Community by Csaba Ajtony (@pits2022).</description>
    <link>https://dev.to/pits2022</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4146619%2Fd0189284-70e3-4ee2-9365-fa0dcac1d0d4.jpg</url>
      <title>DEV Community: Csaba Ajtony</title>
      <link>https://dev.to/pits2022</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/pits2022"/>
    <language>en</language>
    <item>
      <title>Kubernetes Deployments, DaemonSets, and StatefulSets: a Deep Dive from a Production Outage</title>
      <dc:creator>Csaba Ajtony</dc:creator>
      <pubDate>Mon, 28 Sep 2026 08:35:32 +0000</pubDate>
      <link>https://dev.to/pits2022/kubernetes-deployments-daemonsets-and-statefulsets-a-deep-dive-from-a-production-outage-4ghc</link>
      <guid>https://dev.to/pits2022/kubernetes-deployments-daemonsets-and-statefulsets-a-deep-dive-from-a-production-outage-4ghc</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published on the &lt;a href="https://www.professional-it-services.com/kubernetes-deployments-daemonsets-and-statefulsets-a-deep-dive/" rel="noopener noreferrer"&gt;Professional IT Services blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Every introduction to Kubernetes workload controllers gives the same three-line answer: Deployments for stateless apps, DaemonSets for one pod per node, StatefulSets for databases. That answer is correct, and it is also the part that never causes an outage. What causes outages is how each controller behaves when a node is drained, a volume refuses to attach, or a rollout hits a pod that never becomes healthy.&lt;/p&gt;

&lt;p&gt;This deep dive answers the question from a production cluster rather than from the documentation. The worked example is our own &lt;strong&gt;K3s cluster on Hetzner Cloud: 5 nodes running 50 Deployments, 6 DaemonSets and 12 StatefulSets, with 26 persistent volumes&lt;/strong&gt;. And the story that ties it together is the &lt;strong&gt;three-hour outage of 17 June 2026&lt;/strong&gt;, when a routine node upgrade took every WordPress site on the cluster offline, and the workload controllers decided both how bad it got and how it was repaired.&lt;/p&gt;

&lt;p&gt;One boundary first: this article is about &lt;strong&gt;which controller, and how it behaves when things break&lt;/strong&gt;. Sizing — requests, limits, what the Galera pods and the DaemonSets actually consume — is covered separately in &lt;a href="https://www.professional-it-services.com/how-we-cut-kubernetes-resource-overhead-by-50-using-only-built-in-tools/" rel="noopener noreferrer"&gt;how we cut Kubernetes resource overhead by 50% using only built-in tools&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 30-second decision tree
&lt;/h2&gt;

&lt;p&gt;Ask three questions, in this order, and stop at the first yes.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Does the pod work with the node itself&lt;/strong&gt; — its logs, its metrics, its disks, its network, or traffic that arrives at that node? → &lt;strong&gt;DaemonSet.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Does each replica need its own identity and its own data&lt;/strong&gt; that must follow it when it is rescheduled — a database member, a message broker, a quorum peer? → &lt;strong&gt;StatefulSet.&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Everything else&lt;/strong&gt; → &lt;strong&gt;Deployment.&lt;/strong&gt; That includes an app that mounts a single persistent volume — but then choose the update strategy on purpose (see the next section).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In one sentence: &lt;strong&gt;a Deployment is for interchangeable pods, a StatefulSet is for pods that are not interchangeable, and a DaemonSet is for pods that belong to a node rather than to the application.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Deployments: the default, and the persistent-volume trap
&lt;/h2&gt;

&lt;p&gt;A Deployment keeps a given number of identical pods running through a ReplicaSet. Pods get random names, any one can replace any other, and updates roll out gradually, with rollback built in. That is exactly right for web servers, APIs, workers and most of what a cluster runs, which is why &lt;strong&gt;50 of our 68 workload controllers are Deployments&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can a Deployment use persistent storage?
&lt;/h3&gt;

&lt;p&gt;Yes. The previous version of this article said Deployments have ephemeral storage, and that was wrong: a Deployment can mount a PersistentVolumeClaim like any other pod. The trap is not the storage. &lt;strong&gt;It is the default update strategy combined with a ReadWriteOnce volume.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A Deployment's default strategy is &lt;code&gt;RollingUpdate&lt;/code&gt; with &lt;code&gt;maxSurge: 25%&lt;/code&gt; and &lt;code&gt;maxUnavailable: 25%&lt;/code&gt;. Kubernetes rounds the surge &lt;strong&gt;up&lt;/strong&gt; and the unavailability &lt;strong&gt;down&lt;/strong&gt;. With &lt;code&gt;replicas: 1&lt;/code&gt;, that means a surge of 1 and an unavailability of 0: &lt;strong&gt;the new pod starts before the old one stops&lt;/strong&gt;. If the scheduler places the new pod on a different node, the ReadWriteOnce volume is still attached to the old node, the new pod sits in &lt;code&gt;ContainerCreating&lt;/code&gt; with a &lt;code&gt;Multi-Attach error&lt;/code&gt;, and the rollout stalls.&lt;/p&gt;

&lt;p&gt;On our cluster, &lt;strong&gt;15 Deployments mount a PVC. All 15 are ReadWriteOnce Hetzner volumes, and all 15 run a single replica.&lt;/strong&gt; They include the WordPress sites, a mail stack whose &lt;strong&gt;five components share one 100 GiB volume&lt;/strong&gt;, Nextcloud, Grafana and a newsletter server. Every one of them is a candidate for the trap.&lt;/p&gt;

&lt;p&gt;After the June outage we reviewed all fifteen. The rule the review settled on: &lt;strong&gt;&lt;code&gt;Recreate&lt;/code&gt; where it is necessary, &lt;code&gt;RollingUpdate&lt;/code&gt; where it is feasible.&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;Recreate&lt;/code&gt;&lt;/strong&gt; stops the old pod before starting the new one. Each rollout costs a few seconds of downtime, but it cannot deadlock on a volume. Three of the fifteen use it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;RollingUpdate&lt;/code&gt; on one node.&lt;/strong&gt; ReadWriteOnce means &lt;em&gt;one node at a time&lt;/em&gt;, not one pod — the per-pod mode is &lt;code&gt;ReadWriteOncePod&lt;/code&gt;. Two pods on the &lt;strong&gt;same&lt;/strong&gt; node can both mount the volume. So the WordPress sites keep their zero-downtime rollouts, and are kept on one node instead.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Our web chart does it with a &lt;code&gt;pod-group&lt;/code&gt; label and a pod-affinity rule, so the new pod is drawn to the node where its predecessor, and any other pod sharing the volume, already runs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;strategy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;RollingUpdate&lt;/span&gt;
&lt;span class="na"&gt;template&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;pod-group&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{{&lt;/span&gt; &lt;span class="nv"&gt;.Values.podAffinity.group&lt;/span&gt; &lt;span class="pi"&gt;}}&lt;/span&gt;
  &lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;affinity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;podAffinity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;preferredDuringSchedulingIgnoredDuringExecution&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;weight&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;100&lt;/span&gt;
            &lt;span class="na"&gt;podAffinityTerm&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="na"&gt;labelSelector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                &lt;span class="na"&gt;matchExpressions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pod-group&lt;/span&gt;
                    &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;In&lt;/span&gt;
                    &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="pi"&gt;{{&lt;/span&gt; &lt;span class="nv"&gt;.Values.podAffinity.group&lt;/span&gt; &lt;span class="pi"&gt;}}&lt;/span&gt;
              &lt;span class="na"&gt;topologyKey&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kubernetes.io/hostname&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The mail stack and Nextcloud go further and are pinned to a node with a &lt;code&gt;nodeSelector&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you copy this pattern, know its limit.&lt;/strong&gt; &lt;code&gt;preferred&lt;/code&gt; affinity is a hint, not a guarantee: when the node is under pressure, the scheduler is allowed to place the new pod elsewhere, and then you are back to the Multi-Attach error. &lt;code&gt;required&lt;/code&gt; affinity or a &lt;code&gt;nodeSelector&lt;/code&gt; gives the guarantee, at the price of a pod that stays &lt;code&gt;Pending&lt;/code&gt; if that one node is full. If you have not thought about this trade-off for a particular app, &lt;code&gt;Recreate&lt;/code&gt; is the safe default.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does a DaemonSet really run on every node?
&lt;/h2&gt;

&lt;p&gt;No. &lt;strong&gt;A DaemonSet runs one pod on every node it is allowed onto&lt;/strong&gt;, and the difference is where the interesting decisions are.&lt;/p&gt;

&lt;p&gt;Our cluster has a dedicated database node, tainted &lt;code&gt;dedicated=mariadb:NoSchedule&lt;/code&gt; so that only the database lands there. This is what the six DaemonSets do with it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;DaemonSet&lt;/th&gt;
&lt;th&gt;Pods&lt;/th&gt;
&lt;th&gt;Runs on the DB node?&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;fluent-bit&lt;/code&gt; (logs)&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;Yes, explicit toleration&lt;/td&gt;
&lt;td&gt;No logs from the DB node is a blind spot&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;loki-canary&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;Yes, explicit toleration&lt;/td&gt;
&lt;td&gt;Same reason&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;prometheus-node-exporter&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;Yes, tolerates every &lt;code&gt;NoSchedule&lt;/code&gt; taint&lt;/td&gt;
&lt;td&gt;No metrics from the DB node is a blind spot&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;hcloud-csi-node&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;5/5&lt;/td&gt;
&lt;td&gt;Yes, tolerates every taint&lt;/td&gt;
&lt;td&gt;Without the CSI node plugin, no volume attaches on that node&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ingress-nginx-controller&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;4/5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No, by design&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The DB node serves no web traffic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;svclb-ingress-nginx-controller&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;4/5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;No&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Follows the ingress controller&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Same cluster, same taint, and the opposite answer is right for different DaemonSets. &lt;strong&gt;The observability agents must reach the tainted node&lt;/strong&gt; — the database node is exactly where you want logs and metrics when something goes wrong. &lt;strong&gt;The ingress controller must not.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The detail that catches people: the DaemonSet controller automatically adds tolerations for the &lt;strong&gt;built-in&lt;/strong&gt; node-condition taints (&lt;code&gt;not-ready&lt;/code&gt;, &lt;code&gt;unreachable&lt;/code&gt;, &lt;code&gt;memory-pressure&lt;/code&gt;, &lt;code&gt;disk-pressure&lt;/code&gt;, &lt;code&gt;pid-pressure&lt;/code&gt;, &lt;code&gt;unschedulable&lt;/code&gt;, and &lt;code&gt;network-unavailable&lt;/code&gt; for host-network pods). It does &lt;strong&gt;not&lt;/strong&gt; add tolerations for your own taints. Add a custom taint to a node, and every DaemonSet without a matching toleration silently stops covering it. There is no error, only a &lt;code&gt;DESIRED&lt;/code&gt; count one lower than your node count. After tainting a node, &lt;code&gt;kubectl get ds -A&lt;/code&gt; is the check.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why run the ingress controller as a DaemonSet?
&lt;/h3&gt;

&lt;p&gt;Because of how traffic reaches the cluster. A Hetzner load balancer forwards incoming traffic to all worker nodes, so every worker node needs an ingress controller to answer it. The chart supports both modes, and ours is set to DaemonSet:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;controller&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="c1"&gt;# -- Use a `DaemonSet` or `Deployment`&lt;/span&gt;
  &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;DaemonSet&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The old version of this article also claimed that DaemonSets have "no built-in rolling update strategy (until Kubernetes 1.7+)". That is long obsolete: &lt;code&gt;RollingUpdate&lt;/code&gt; is the default in the &lt;code&gt;apps/v1&lt;/code&gt; API, and all six of our DaemonSets use it with &lt;code&gt;maxUnavailable: 1&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  StatefulSets as they actually run in production
&lt;/h2&gt;

&lt;p&gt;The textbook StatefulSet gives each pod a stable ordinal name (&lt;code&gt;db-0&lt;/code&gt;, &lt;code&gt;db-1&lt;/code&gt;, &lt;code&gt;db-2&lt;/code&gt;), a stable DNS entry through a headless Service, and its own PersistentVolumeClaim from a &lt;code&gt;volumeClaimTemplate&lt;/code&gt;, which follows the pod when it is rescheduled. By default it starts pods one at a time in ordinal order (&lt;code&gt;podManagementPolicy: OrderedReady&lt;/code&gt;) and rolls updates in reverse order, with a &lt;code&gt;partition&lt;/code&gt; to stage them.&lt;/p&gt;

&lt;p&gt;Our production database is a &lt;strong&gt;three-node MariaDB Galera cluster&lt;/strong&gt;, and it shows how differently a real one is configured. It is not a hand-written StatefulSet. It is declared as a &lt;code&gt;MariaDB&lt;/code&gt; custom resource, and &lt;strong&gt;mariadb-operator&lt;/strong&gt; generates the StatefulSet from it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;k8s.mariadb.com/v1alpha1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;MariaDB&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;replicas&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;3&lt;/span&gt;
  &lt;span class="na"&gt;galera&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;enabled&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="na"&gt;affinity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;podAntiAffinity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;requiredDuringSchedulingIgnoredDuringExecution&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;labelSelector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="na"&gt;matchExpressions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
              &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;app.kubernetes.io/instance&lt;/span&gt;
                &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;In&lt;/span&gt;
                &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
                  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;mariadb-galera-new&lt;/span&gt;
          &lt;span class="na"&gt;topologyKey&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kubernetes.io/hostname&lt;/span&gt;
  &lt;span class="na"&gt;podDisruptionBudget&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;minAvailable&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The ordering people choose StatefulSets for is switched off
&lt;/h3&gt;

&lt;p&gt;The generated StatefulSet runs with &lt;strong&gt;&lt;code&gt;podManagementPolicy: Parallel&lt;/code&gt;&lt;/strong&gt; and &lt;strong&gt;&lt;code&gt;updateStrategy: OnDelete&lt;/code&gt;&lt;/strong&gt;. The ordered, one-at-a-time behaviour that most articles list as the main reason to use a StatefulSet is turned off — and the operator does the ordering itself, rolling the &lt;strong&gt;replicas first and the primary last&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The reason is how Galera works. A Galera member that cannot see its peers cannot form a quorum, so "start pod-0, wait until it is Ready, then start pod-1" can stall a full-cluster restart. The operator starts the members together, and applies the order that actually matters for a database — never restart the node taking writes until the others are updated — in its own logic.&lt;/p&gt;

&lt;p&gt;The rest of the configuration is where the operational safety lives:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Required anti-affinity on the hostname&lt;/strong&gt;, so no two members share a node, plus a &lt;code&gt;nodeSelector&lt;/code&gt; and a toleration that admit them to the database nodes only.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A PodDisruptionBudget of &lt;code&gt;minAvailable: 2&lt;/code&gt;.&lt;/strong&gt; A three-node Galera needs two nodes for quorum, so a node drain may take down at most one member at a time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;persistentVolumeClaimRetentionPolicy: Retain&lt;/code&gt; for both deletion and scale-down.&lt;/strong&gt; Deleting the StatefulSet or scaling it down never deletes the data. In June this was what made the repair possible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Two &lt;code&gt;volumeClaimTemplates&lt;/code&gt; per pod&lt;/strong&gt;: a 20 GiB data volume and a 100 MiB Galera state volume. The 100 MiB claims are bound at &lt;strong&gt;10 GiB&lt;/strong&gt;, because Hetzner's minimum volume size is 10 GB — three times over, for about 300 MB of state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Volumes are tied to a location.&lt;/strong&gt; A Hetzner volume attaches only to a server in the same location. A PVC created in one datacenter cannot follow its pod to another one; moving it means copying the data into a new PVC.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  17 June 2026: what the controllers did during a three-hour outage
&lt;/h2&gt;

&lt;p&gt;The Galera cluster in production at the time was the previous one: a Bitnami chart StatefulSet with the default ordered pod management. This is how a routine maintenance task became a total outage.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A routine worker-node upgrade&lt;/strong&gt; drained a node.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Galera lost its Primary component.&lt;/strong&gt; All three members went non-Primary, and &lt;strong&gt;every WordPress site on the cluster returned "Error establishing a database connection"&lt;/strong&gt;. Contributing factors, as the runbook records them: one member was running in a different datacenter, the pods sat at about &lt;strong&gt;99% of a 5 GiB memory limit&lt;/strong&gt;, and heavy database dumps were running at the same time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The CSI pods could not be scheduled either.&lt;/strong&gt; No CSI, no volume attachments — and from that point everything that needed a persistent volume was down, not just the database.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recovery started by bootstrapping node-0&lt;/strong&gt; with an empty cluster address (&lt;code&gt;gcomm://&lt;/code&gt;), so it formed a new Primary component on its own.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The StatefulSet's rolling update was stuck.&lt;/strong&gt; Under &lt;code&gt;OrderedReady&lt;/code&gt;, the controller will not roll past a pod that never becomes Ready; the Kubernetes documentation says that in this state you have to delete the broken pods by hand. The fix was to delete the StatefulSet &lt;strong&gt;without deleting its pods or PVCs&lt;/strong&gt; and recreate it:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   kubectl delete statefulset &amp;lt;name&amp;gt; &lt;span class="nt"&gt;--cascade&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;orphan
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Three-node HA came back through a staged rollout with &lt;code&gt;partition=1&lt;/code&gt;.&lt;/strong&gt; Node-0 kept serving untouched, nodes 1 and 2 rejoined through a full state transfer (SST), and node-0 rolled last. The memory limit went from 5 GiB to 7 GiB.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Total downtime: &lt;strong&gt;about three hours. No data was lost.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  What changed afterwards
&lt;/h3&gt;

&lt;p&gt;The Bitnami image in use had been moved to the frozen &lt;code&gt;bitnamilegacy&lt;/code&gt; repository, with no further security updates, so the cluster was moved to mariadb-operator. The first cutover attempt — a full dump and reload under a write freeze — took &lt;strong&gt;about 60 minutes&lt;/strong&gt; on a three-node Galera, which is far too long to freeze writes. It was redone the next day as &lt;strong&gt;asynchronous GTID replication&lt;/strong&gt; from the old cluster to the new one, followed by a write freeze of &lt;strong&gt;seconds&lt;/strong&gt; and a switch of the Service selector.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The database's DNS name never changed.&lt;/strong&gt; The old &lt;code&gt;mariadb-galera&lt;/code&gt; Service now selects the new primary, so not a single tenant configuration file had to be touched. At the same time, the database and several application volumes were moved from the old datacenter to the main one — by copying the data into new PVCs, since a volume cannot move in place.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lessons that apply to any cluster
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A node drain is decided by your PodDisruptionBudgets, taints and tolerations&lt;/strong&gt; — check what each controller will do before upgrading nodes, including the CSI components.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Know what ordered pod management does with a pod that never becomes Ready.&lt;/strong&gt; It waits. The rollout will not repair itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;--cascade=orphan&lt;/code&gt; plus a &lt;code&gt;Retain&lt;/code&gt; PVC policy&lt;/strong&gt; lets you rebuild a controller without touching the running pods or the data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the Service name stable.&lt;/strong&gt; The DNS name is the contract with every client of the database; everything behind it can be replaced.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Move data safely:&lt;/strong&gt; never dump inside the database pod (it will run out of memory), never stream large data through the kube-apiserver (a small control-plane node will fall over), copy pod to pod inside the cluster with a bandwidth cap (&lt;code&gt;pv -L 8m&lt;/code&gt;), one database at a time, with &lt;code&gt;set -o pipefail&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Key differences at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Deployment&lt;/th&gt;
&lt;th&gt;DaemonSet&lt;/th&gt;
&lt;th&gt;StatefulSet&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Manages&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Interchangeable replicas&lt;/td&gt;
&lt;td&gt;One pod per eligible node&lt;/td&gt;
&lt;td&gt;Replicas with stable identity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Pod names&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Random hash&lt;/td&gt;
&lt;td&gt;Random hash, one per node&lt;/td&gt;
&lt;td&gt;Ordinal: &lt;code&gt;name-0&lt;/code&gt;, &lt;code&gt;name-1&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Placement&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Scheduler decides&lt;/td&gt;
&lt;td&gt;Every node it tolerates&lt;/td&gt;
&lt;td&gt;Scheduler decides; anti-affinity usually required&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Storage&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Any PVC; shared RWO volumes need care&lt;/td&gt;
&lt;td&gt;Usually node-local paths&lt;/td&gt;
&lt;td&gt;Own PVC per pod via &lt;code&gt;volumeClaimTemplates&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Default update&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;RollingUpdate&lt;/code&gt;, 25% surge / 25% unavailable&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;RollingUpdate&lt;/code&gt;, &lt;code&gt;maxUnavailable: 1&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;RollingUpdate&lt;/code&gt; in reverse ordinal order, with &lt;code&gt;partition&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Ordering&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;OrderedReady&lt;/code&gt; by default; clustered DB operators often use &lt;code&gt;Parallel&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Typical failure&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Multi-Attach error on a single-replica RWO rollout&lt;/td&gt;
&lt;td&gt;Silently skips nodes with custom taints&lt;/td&gt;
&lt;td&gt;Rollout stuck on a pod that never becomes Ready&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Use for&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Web apps, APIs, workers&lt;/td&gt;
&lt;td&gt;Logs, metrics, CSI, ingress on every node&lt;/td&gt;
&lt;td&gt;Databases, brokers, quorum-based systems&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Final thoughts
&lt;/h2&gt;

&lt;p&gt;Choosing the controller takes thirty seconds with the decision tree above. The work is in the behaviour around it: the update strategy of a Deployment with a volume, the tolerations of every DaemonSet, and the pod management, disruption budget and retention policy of every StatefulSet. On our cluster, each of those was either a cause of the June outage or part of its repair.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Keeping the resulting cluster efficient is the other half: our production K3s audit — &lt;a href="https://www.professional-it-services.com/how-we-cut-kubernetes-resource-overhead-by-50-using-only-built-in-tools/" rel="noopener noreferrer"&gt;how we cut Kubernetes resource overhead by 50% using only built-in tools&lt;/a&gt; — shows what over-requested StatefulSets and cluster-wide DaemonSets actually cost once they are measured.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>devops</category>
      <category>k3s</category>
      <category>database</category>
    </item>
  </channel>
</rss>
