<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Siddharth More</title>
    <description>The latest articles on DEV Community by Siddharth More (@siddharthajmore).</description>
    <link>https://dev.to/siddharthajmore</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3276428%2Fbc3b2dc8-6d71-4060-9c3e-c7d0e12eae09.jpg</url>
      <title>DEV Community: Siddharth More</title>
      <link>https://dev.to/siddharthajmore</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/siddharthajmore"/>
    <language>en</language>
    <item>
      <title>How Kubernetes Nodes Learn to Say No: Taints and Tolerations</title>
      <dc:creator>Siddharth More</dc:creator>
      <pubDate>Tue, 22 Sep 2026 04:58:57 +0000</pubDate>
      <link>https://dev.to/siddharthajmore/when-nodes-start-rejecting-pods-taints-and-tolerations-1kf3</link>
      <guid>https://dev.to/siddharthajmore/when-nodes-start-rejecting-pods-taints-and-tolerations-1kf3</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Taints go on nodes, tolerations go on pods. A tainted node rejects pods by default, and a toleration is what lets a specific pod back in. It is opposite of node selectors and affinity, where the pod was the one doing the choosing.&lt;/li&gt;
&lt;li&gt;Three effects: &lt;code&gt;NoSchedule&lt;/code&gt;, &lt;code&gt;PreferNoSchedule&lt;/code&gt;, &lt;code&gt;NoExecute&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;NoExecute&lt;/code&gt; is the odd one out. It can evict pods that are already running, not just block new ones from landing.&lt;/li&gt;
&lt;li&gt;Tolerating a taint doesn't mean a pod prefers that node. It just means the taint won't stop it from scheduling on the node.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Skip around if you already know the basics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The problem this actually solves&lt;/li&gt;
&lt;li&gt;Taints, tolerations, and the three effects&lt;/li&gt;
&lt;li&gt;Where this actually shows up&lt;/li&gt;
&lt;li&gt;What the scheduler's doing under the hood&lt;/li&gt;
&lt;li&gt;Taints and tolerations vs node affinity&lt;/li&gt;
&lt;li&gt;What goes wrong, and why&lt;/li&gt;
&lt;li&gt;Try it yourself&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  WHY: the problem
&lt;/h2&gt;

&lt;p&gt;Node selectors and affinity, from Part 2, are opt-in. A pod decides which nodes it wants, and the scheduler tries to honor that.&lt;/p&gt;

&lt;p&gt;Sometimes you need the opposite. A node that rejects almost everything by default, unless a pod specifically says it's fine with that. Control-plane nodes are the classic example, you don't want regular application pods landing there and competing with the control plane for resources.&lt;/p&gt;

&lt;p&gt;That's what taints and tolerations are for. Instead of a pod pulling toward a node, a node pushes pods away, and only pods that explicitly tolerate that push can get scheduled on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  WHAT: taints, tolerations, and effects
&lt;/h2&gt;

&lt;p&gt;A taint lives on a node. It's a key, a value, and an effect, written as &lt;code&gt;key=value:effect&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A toleration lives on a pod. It lists the key, value, and effect the pod is willing to tolerate. If it matches, the taint doesn't block that pod.&lt;/p&gt;

&lt;p&gt;Here's a way to picture it: a taint is bug repellent sprayed on a node. A toleration is a bug that's immune to that specific spray.&lt;/p&gt;

&lt;p&gt;The three effects behave differently, and the analogy maps onto all of them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;NoSchedule&lt;/code&gt;&lt;/strong&gt; - full-strength spray. Immune bugs land fine, everything else won't come near. Bugs that landed before the spray went down aren't touched either way, this only stops new arrivals.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;PreferNoSchedule&lt;/code&gt;&lt;/strong&gt; - a lighter dose. Bugs would rather avoid it, but if every other spot's taken, they'll land there anyway. Same as above, anything that landed earlier just stays put.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;NoExecute&lt;/code&gt;&lt;/strong&gt; - once the spray's applied, only immune bugs can land or stay. Every other bug gets rejected, whether it landed before the spray or was about to land after.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Setting a taint on a node:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl taint nodes node-1 &lt;span class="nv"&gt;dedicated&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;gpu:NoSchedule
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Tolerating it from a pod:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Pod&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gpu-pod&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;tolerations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dedicated"&lt;/span&gt;
    &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Equal"&lt;/span&gt;
    &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpu"&lt;/span&gt;
    &lt;span class="na"&gt;effect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;NoSchedule"&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;operator&lt;/code&gt; can also be &lt;code&gt;Exists&lt;/code&gt; instead of &lt;code&gt;Equal&lt;/code&gt;, if a pod should tolerate a key regardless of its value:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;tolerations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dedicated"&lt;/span&gt;
  &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Exists"&lt;/span&gt;
  &lt;span class="na"&gt;effect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;NoSchedule"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A node isn't limited to one taint. When there are several taints, a pod has to tolerate every &lt;code&gt;NoSchedule&lt;/code&gt; and &lt;code&gt;NoExecute&lt;/code&gt; taint on that node to be allowed there, each one gets checked independently, and missing even one will keep the pod out. &lt;code&gt;PreferNoSchedule&lt;/code&gt; taints don't count toward that requirement, they only ever affect scoring, never a hard block.&lt;/p&gt;

&lt;p&gt;One more piece worth knowing: &lt;code&gt;tolerationSeconds&lt;/code&gt;. It only applies to &lt;code&gt;NoExecute&lt;/code&gt;, and it changes what the toleration itself means. Instead of tolerating the taint forever, it's saying "I tolerate this taint, but only for N seconds." The countdown starts the moment the taint shows up. Until it runs out, the pod counts as tolerating the taint and stays put. Once it expires, Kubernetes treats the pod as if it no longer tolerates that taint at all, and evicts it, exactly like a pod that never had a matching toleration in the first place.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;tolerations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;dedicated"&lt;/span&gt;
  &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Equal"&lt;/span&gt;
  &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpu"&lt;/span&gt;
  &lt;span class="na"&gt;effect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;NoExecute"&lt;/span&gt;
  &lt;span class="na"&gt;tolerationSeconds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;3600&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;📖 Docs: &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/taint-and-toleration/" rel="noopener noreferrer"&gt;Taints and tolerations&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  WHEN: where this actually shows up
&lt;/h2&gt;

&lt;p&gt;A few real situations where taints and tolerations show up:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Control-plane nodes, tainted by default so regular workloads stay off them.&lt;/li&gt;
&lt;li&gt;Dedicated GPU or other specialized-hardware nodes, so only the workloads that need them land there.&lt;/li&gt;
&lt;li&gt;Kubernetes itself uses this mechanism internally. When a node goes unreachable or stops responding, the node controller taints it automatically.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  HOW: what's actually happening
&lt;/h2&gt;

&lt;p&gt;Taints get checked during the same filtering phase from Part 1, just inverted. Normally filtering asks "can this node run the pod?" Here it's more like "has this node been told to reject the pod?" If yes and if there's no matching toleration on pod, the node gets filtered out, same as any other hard requirement.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;PreferNoSchedule&lt;/code&gt; doesn't touch filtering at all. It works through scoring, same as preferred node affinity in Part 2, nudging the scheduler away from tainted nodes without ruling them out.&lt;/p&gt;

&lt;p&gt;What happens on a node that's already running pods matters here, (I am intentionally repeating these points to make it stick to your memory) and the three effects don't behave the same way:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Taint a live node with &lt;code&gt;NoSchedule&lt;/code&gt; or &lt;code&gt;PreferNoSchedule&lt;/code&gt;, and nothing happens to pods already running there. These only affect future scheduling decisions, not the pods sitting there right now.&lt;/li&gt;
&lt;li&gt;Taint it with &lt;code&gt;NoExecute&lt;/code&gt;, and any pod without a matching toleration gets evicted, immediately, or after the grace period if &lt;code&gt;tolerationSeconds&lt;/code&gt; is set. This is the one real exception to "scheduling is a one-time decision" from Part 1, worth closing that loop here.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The automatic tainting mentioned in previous section works off a fixed set of built-in keys, all under the &lt;code&gt;node.kubernetes.io/&lt;/code&gt; prefix: &lt;code&gt;not-ready&lt;/code&gt;, &lt;code&gt;unreachable&lt;/code&gt;, &lt;code&gt;memory-pressure&lt;/code&gt;, &lt;code&gt;disk-pressure&lt;/code&gt;, &lt;code&gt;pid-pressure&lt;/code&gt;, &lt;code&gt;network-unavailable&lt;/code&gt;, and &lt;code&gt;unschedulable&lt;/code&gt;. The first two carry &lt;code&gt;NoExecute&lt;/code&gt;, which is what actually triggers eviction the moment a node's condition flips.&lt;/p&gt;

&lt;p&gt;That sounds like it should mean every pod in the cluster gets evicted the instant a node blips offline for a second, but Kubernetes doesn't let that happen by default. Every pod automatically gets a built-in toleration for &lt;code&gt;not-ready&lt;/code&gt; and &lt;code&gt;unreachable&lt;/code&gt;, capped at 300 seconds, whether or not one was written by hand. That's the buffer that keeps a brief network hiccup from turning into a mass eviction.&lt;/p&gt;

&lt;h2&gt;
  
  
  DIFFERENCE: taints and tolerations vs node affinity
&lt;/h2&gt;

&lt;p&gt;Node affinity and taints solve a similar-sounding problem from opposite directions.&lt;/p&gt;

&lt;p&gt;Affinity is a pod pulling toward a node it wants. Taints are a node pushing pods away by default, and a toleration is what lets a specific pod back in.&lt;/p&gt;

&lt;p&gt;Here's where that difference actually shows up. Back in Part 2, node affinity let a pod prefer or require a specific node, but that was only the pod's side of the deal. Nothing stopped some other pod, one with no affinity rule at all, from landing on that same node too. Affinity never restricts the node itself.&lt;/p&gt;

&lt;p&gt;Taints and tolerations exist for exactly that gap. If the goal is "&lt;em&gt;no pod should schedule here unless it explicitly tolerates this&lt;/em&gt;" affinity can't do that on its own, that's what a taint is for.&lt;/p&gt;

&lt;p&gt;Tolerating a taint doesn't mean a pod prefers that node either, or that it'll get scheduled there. It just means the taint won't block it. &lt;strong&gt;If the goal is a pod landing on GPU nodes specifically, not just being allowed there, tolerations usually get paired with node affinity on top.&lt;/strong&gt; The taint keeps everyone else out, the affinity pulls the right pod in.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAILURE: what goes wrong
&lt;/h2&gt;

&lt;p&gt;The tolerate-doesn't-mean-prefer mix-up is the most common one. A pod with the right toleration can still land on an untainted node it has nothing to do with, because a toleration only removes a restriction, it doesn't express a preference.&lt;/p&gt;

&lt;p&gt;The second one is more of a hazard than a misconception. Adding a &lt;code&gt;NoExecute&lt;/code&gt; taint to a node that's already serving traffic can evict running pods with no warning, if that wasn't accounted for. Worth checking what's already running and what it tolerates before tainting a node people are relying on.&lt;/p&gt;

&lt;p&gt;One more thing worth flagging now and covering properly later: DaemonSet pods tolerate several of these built-in taints automatically, which is part of why they keep running in places regular pods can't. More on that in Part 5.&lt;/p&gt;

&lt;h2&gt;
  
  
  PRACTICE: taint, tolerate, check
&lt;/h2&gt;

&lt;p&gt;Taint a node:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl taint nodes node-1 &lt;span class="nv"&gt;dedicated&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;gpu:NoSchedule
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Deploy the toleration pod from the WHAT section, plus a plain one with no toleration at all:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Pod&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;plain-pod&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl apply &lt;span class="nt"&gt;-f&lt;/span&gt; gpu-pod.yaml
kubectl apply &lt;span class="nt"&gt;-f&lt;/span&gt; plain-pod.yaml
kubectl get pods &lt;span class="nt"&gt;-o&lt;/span&gt; wide
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;plain-pod&lt;/code&gt; should avoid &lt;code&gt;node-1&lt;/code&gt; entirely. Confirm the taint is actually there:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl describe node node-1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look for a &lt;code&gt;Taints:&lt;/code&gt; line near the top of the output.&lt;/p&gt;

&lt;p&gt;To remove the taint later:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl taint nodes node-1 &lt;span class="nv"&gt;dedicated&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;gpu:NoSchedule-
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That trailing &lt;code&gt;-&lt;/code&gt; is easy to miss and just as easy to forget, so it's worth calling out on its own.&lt;/p&gt;

&lt;p&gt;Next up, Part 4: pod affinity and anti-affinity, where instead of a pod caring about a node, it starts caring about other pods.&lt;/p&gt;

&lt;p&gt;If anything here didn't land, or you'd explain it differently, drop a comment. Doubts and pushback are both welcome.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>devops</category>
      <category>beginners</category>
      <category>cloudnative</category>
    </item>
    <item>
      <title>How does Kubernetes decide where GPU workloads or SSD-heavy databases should run?

Published Part 2 of my Kubernetes scheduling series. This one covers node selectors and node affinity.</title>
      <dc:creator>Siddharth More</dc:creator>
      <pubDate>Fri, 18 Sep 2026 10:57:35 +0000</pubDate>
      <link>https://dev.to/siddharthajmore/how-does-kubernetes-decide-where-gpu-workloads-or-ssd-heavy-databases-should-run-published-part-5g16</link>
      <guid>https://dev.to/siddharthajmore/how-does-kubernetes-decide-where-gpu-workloads-or-ssd-heavy-databases-should-run-published-part-5g16</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/siddharthajmore/kubernetes-scheduling-101-what-really-happens-before-your-pod-runs-3khk" class="crayons-story__hidden-navigation-link"&gt;Kubernetes Scheduling 101: What Really Happens Before Your Pod Runs&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/siddharthajmore" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3276428%2Fbc3b2dc8-6d71-4060-9c3e-c7d0e12eae09.jpg" alt="siddharthajmore profile" class="crayons-avatar__image" width="800" height="800"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/siddharthajmore" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Siddharth More
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Siddharth More
                
                
              
              &lt;div id="story-author-preview-content-4654537" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/siddharthajmore" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3276428%2Fbc3b2dc8-6d71-4060-9c3e-c7d0e12eae09.jpg" class="crayons-avatar__image" alt="" width="800" height="800"&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Siddharth More&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/siddharthajmore/kubernetes-scheduling-101-what-really-happens-before-your-pod-runs-3khk" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 15&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/siddharthajmore/kubernetes-scheduling-101-what-really-happens-before-your-pod-runs-3khk" id="article-link-4654537"&gt;
          Kubernetes Scheduling 101: What Really Happens Before Your Pod Runs
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/kubernetes"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;kubernetes&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/beginners"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;beginners&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/devops"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;devops&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/siddharthajmore/kubernetes-scheduling-101-what-really-happens-before-your-pod-runs-3khk" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="24" height="24"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;1&lt;span class="hidden s:inline"&gt;&amp;nbsp;reaction&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/siddharthajmore/kubernetes-scheduling-101-what-really-happens-before-your-pod-runs-3khk#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              2&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            4 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
      <category>devops</category>
      <category>infrastructure</category>
      <category>kubernetes</category>
    </item>
    <item>
      <title>Telling the Kube-Scheduler Where Pods Can Run: Node Selectors and Node Affinity</title>
      <dc:creator>Siddharth More</dc:creator>
      <pubDate>Fri, 18 Sep 2026 06:25:26 +0000</pubDate>
      <link>https://dev.to/siddharthajmore/node-selectors-and-node-affinity-telling-the-scheduler-where-pods-can-run-1loj</link>
      <guid>https://dev.to/siddharthajmore/node-selectors-and-node-affinity-telling-the-scheduler-where-pods-can-run-1loj</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;nodeSelector&lt;/code&gt; matches a pod to a node using labels. Simple, exact match, nothing fancier.&lt;/li&gt;
&lt;li&gt;Node affinity does the same job with more expressive rules: &lt;code&gt;In&lt;/code&gt;, &lt;code&gt;NotIn&lt;/code&gt;, &lt;code&gt;Exists&lt;/code&gt;, &lt;code&gt;DoesNotExist&lt;/code&gt;, &lt;code&gt;Gt&lt;/code&gt;, &lt;code&gt;Lt&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Affinity comes in two flavors: &lt;code&gt;required&lt;/code&gt; (hard rule) and &lt;code&gt;preferred&lt;/code&gt; (soft nudge, weighted).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;IgnoredDuringExecution&lt;/code&gt; means once the pod's running, changes to the node's labels don't affect it. No eviction, no re-check.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Skip around if you already know the basics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The problem this actually solves&lt;/li&gt;
&lt;li&gt;nodeSelector vs node affinity, explained&lt;/li&gt;
&lt;li&gt;Where node selection actually matters&lt;/li&gt;
&lt;li&gt;What the scheduler's doing under the hood&lt;/li&gt;
&lt;li&gt;nodeSelector vs affinity, the trade-offs&lt;/li&gt;
&lt;li&gt;What goes wrong, and why&lt;/li&gt;
&lt;li&gt;Try it yourself&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  WHY: the problem
&lt;/h2&gt;

&lt;p&gt;In Part 1, we saw kube-scheduler filter out nodes that can't run a pod, then score whatever's left. That's default. Most of the time, you don't need to touch it.&lt;/p&gt;

&lt;p&gt;But sometimes the scheduler doesn't know something you know. Maybe only some of your nodes have GPUs. Maybe a database needs SSD-backed storage, and a spinning disk will tank its performance. Nothing in a plain pod spec tells the scheduler to care about that.&lt;/p&gt;

&lt;p&gt;Node selectors and node affinity exist to close that gap. They let you attach the missing context yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  WHAT: the two mechanisms
&lt;/h2&gt;

&lt;p&gt;Both work off the same idea: label your nodes, then tell the pod which labels to look for.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;nodeSelector&lt;/code&gt;&lt;/strong&gt; is the simple version. Give it one or more key-value pairs, and the pod only lands on a node that has every one of them.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Pod&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nodeselector-pod&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;nodeSelector&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;disktype&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ssd&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Node affinity&lt;/strong&gt; does the same job, but with real operators instead of one flat match: &lt;code&gt;In&lt;/code&gt;, &lt;code&gt;NotIn&lt;/code&gt;, &lt;code&gt;Exists&lt;/code&gt;, &lt;code&gt;DoesNotExist&lt;/code&gt;, &lt;code&gt;Gt&lt;/code&gt;, &lt;code&gt;Lt&lt;/code&gt;. &lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Worth noting: there's no separate "node anti-affinity" field - using NotIn or DoesNotExist as the operator is what gets you that behavior, keeping a pod away from nodes that match instead of drawing it toward them.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It also splits into two types:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;requiredDuringSchedulingIgnoredDuringExecution&lt;/code&gt; - a hard rule, functionally like &lt;code&gt;nodeSelector&lt;/code&gt; but more expressive&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;preferredDuringSchedulingIgnoredDuringExecution&lt;/code&gt; - a soft rule with a weight (1-100) that nudges the scheduler's scoring, without blocking the pod if nothing matches
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Pod&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;affinity-pod&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;affinity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;nodeAffinity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;requiredDuringSchedulingIgnoredDuringExecution&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;nodeSelectorTerms&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;matchExpressions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;disktype&lt;/span&gt;
            &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;In&lt;/span&gt;
            &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;hdd&lt;/span&gt;
      &lt;span class="na"&gt;preferredDuringSchedulingIgnoredDuringExecution&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;weight&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;50&lt;/span&gt;
        &lt;span class="na"&gt;preference&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;matchExpressions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;zone&lt;/span&gt;
            &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;In&lt;/span&gt;
            &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;us-east-1a&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read that as: this pod &lt;strong&gt;MUST&lt;/strong&gt; land on an HDD node, and it would prefer &lt;code&gt;us-east-1a&lt;/code&gt;, but it won't sit &lt;code&gt;Pending&lt;/code&gt; for the zone to match.&lt;/p&gt;

&lt;p&gt;One thing that trips people up: &lt;code&gt;nodeSelectorTerms&lt;/code&gt; doesn't just take &lt;code&gt;matchExpressions&lt;/code&gt;. It also takes &lt;code&gt;matchFields&lt;/code&gt;, and the two aren't interchangeable.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;matchExpressions&lt;/code&gt; matches against node &lt;strong&gt;labels&lt;/strong&gt;, the same labels you set earlier. &lt;code&gt;matchFields&lt;/code&gt; matches against node &lt;strong&gt;fields&lt;/strong&gt; instead, things like the node's actual name in the API, not something someone labeled.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;matchFields&lt;/code&gt; lets node affinity match against supported node fields rather than labels. A common example is matching &lt;code&gt;metadata.name&lt;/code&gt; to target a particular node.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Pod&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pinned-to-node-1&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;affinity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;nodeAffinity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;requiredDuringSchedulingIgnoredDuringExecution&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;nodeSelectorTerms&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;matchFields&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;metadata.name&lt;/span&gt;
            &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;In&lt;/span&gt;
            &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;node-1&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Worth being honest about this one: it's rarely the right call. Hard-coding a node name defeats the whole point of a label-based system, and if that node ever gets replaced, your pod just stops scheduling. Reach for &lt;code&gt;matchExpressions&lt;/code&gt; and labels first.&lt;/p&gt;

&lt;p&gt;There's one more thing worth nailing down: how multiple expressions and values actually combine. It's not obvious just from looking at the YAML.&lt;/p&gt;

&lt;p&gt;Inside one &lt;code&gt;matchExpressions&lt;/code&gt; list, every entry has to match. That's an AND.&lt;/p&gt;

&lt;p&gt;Inside one expression's &lt;code&gt;values&lt;/code&gt; list, matching any single value is enough. That's an OR.&lt;/p&gt;

&lt;p&gt;And across separate entries in &lt;code&gt;nodeSelectorTerms&lt;/code&gt; itself, only one term has to fully match. That's another OR, one level up.&lt;/p&gt;

&lt;p&gt;Stacked together, it looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;affinity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;nodeAffinity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;requiredDuringSchedulingIgnoredDuringExecution&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;nodeSelectorTerms&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;matchExpressions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;zone&lt;/span&gt;
            &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;In&lt;/span&gt;
            &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;us-east-1a&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;us-east-1b&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;disktype&lt;/span&gt;
            &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;In&lt;/span&gt;
            &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;ssd&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;matchExpressions&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;zone&lt;/span&gt;
            &lt;span class="na"&gt;operator&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;In&lt;/span&gt;
            &lt;span class="na"&gt;values&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;us-west-2a&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read that as:&lt;br&gt;
&lt;strong&gt;(&lt;/strong&gt;zone is (&lt;code&gt;us-east-1a&lt;/code&gt; &lt;strong&gt;OR&lt;/strong&gt; &lt;code&gt;us-east-1b&lt;/code&gt;), &lt;strong&gt;AND&lt;/strong&gt; disktype is &lt;code&gt;ssd&lt;/code&gt;&lt;strong&gt;) &lt;br&gt;
OR &lt;br&gt;
(&lt;/strong&gt;zone is &lt;code&gt;us-west-2a&lt;/code&gt;&lt;strong&gt;)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It looks like a lot stacked up like that, but it's really just three separate rules, each with its own logic, nested inside each other.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;You can also use both nodeSelector and nodeAffinity,in that case a node must satisfy both before the Pod can be scheduled.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;📖 Docs: &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/#nodeselector" rel="noopener noreferrer"&gt;nodeSelector&lt;/a&gt; · &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/#node-affinity" rel="noopener noreferrer"&gt;node affinity&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  WHEN: where node selection actually matters
&lt;/h2&gt;

&lt;p&gt;A few real situations where controlling Pod placement makes sense, regardless of which mechanism you reach for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GPU workloads that should only run on GPU-equipped nodes&lt;/li&gt;
&lt;li&gt;Databases or caches that need SSD-backed storage&lt;/li&gt;
&lt;li&gt;Workloads you want kept in a specific zone, without hard-blocking if that zone's full&lt;/li&gt;
&lt;li&gt;Compliance or licensing rules that pin certain workloads to certain hardware&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If none of that applies, the scheduler's default behavior is usually fine on its own.&lt;/p&gt;
&lt;h2&gt;
  
  
  HOW: what's actually happening
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;nodeSelector&lt;/code&gt; and the required affinity rule both get evaluated during the filtering phase from Part 1. Any node that doesn't match gets thrown out before scoring even starts.&lt;/p&gt;

&lt;p&gt;The preferred affinity rule works differently. It doesn't touch filtering at all. Once the feasible set is decided, the scheduler adds the rule's weight to a node's score if that node matches. Highest score wins, same as any other scoring factor.&lt;/p&gt;

&lt;p&gt;Both mechanisms share one behavior worth calling out: they're &lt;em&gt;ignored during execution&lt;/em&gt;. The rule only gets checked once, at scheduling time. If someone relabels the node an hour later, nothing happens to the pod already running there.&lt;/p&gt;
&lt;h2&gt;
  
  
  TRADE-OFFS: nodeSelector vs affinity
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;nodeSelector&lt;/code&gt; is less YAML and easier to read. If the requirement really is "must have this one label," there's no reason to reach for affinity instead.&lt;/p&gt;

&lt;p&gt;The moment you need OR logic, negation, or a "prefer but don't require," &lt;code&gt;nodeSelector&lt;/code&gt; can't do it and affinity can. That flexibility costs a bit more nesting, but there's no real downside beyond that.&lt;/p&gt;

&lt;p&gt;One thing neither of these solves: pod-to-pod placement, like "run this next to that" or "spread these apart." That's pod affinity and anti-affinity, coming in Part 4. And if the goal is nodes actively rejecting pods rather than pods choosing nodes, that's taints and tolerations, next up in Part 3.&lt;/p&gt;
&lt;h2&gt;
  
  
  FAILURE: what goes wrong
&lt;/h2&gt;

&lt;p&gt;The most common failure: you set a required rule, no node matches it, and the pod remains &lt;code&gt;Pending&lt;/code&gt; until a suitable node becomes available or the scheduling constraints change. There's no obvious error, just a pod that never starts. &lt;code&gt;kubectl describe pod&lt;/code&gt;, specifically its Events section, is where you'll actually see the scheduling failure reason.&lt;/p&gt;

&lt;p&gt;The second one catches people off guard: because these rules are &lt;code&gt;IgnoredDuringExecution&lt;/code&gt;, relabeling a node does nothing to pods already scheduled there. If you're expecting a label change to trigger a reschedule, it won't. You'd have to delete and recreate the pod yourself.&lt;/p&gt;
&lt;h2&gt;
  
  
  PRACTICE: label, deploy, check
&lt;/h2&gt;

&lt;p&gt;Label two nodes differently:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl label node node-1 &lt;span class="nv"&gt;disktype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;ssd
kubectl label node node-2 &lt;span class="nv"&gt;disktype&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;hdd
kubectl get nodes &lt;span class="nt"&gt;--show-labels&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Save the two pod specs from the WHAT section as &lt;code&gt;nodeselector-pod.yaml&lt;/code&gt; and &lt;code&gt;affinity-pod.yaml&lt;/code&gt;, then deploy both:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl apply &lt;span class="nt"&gt;-f&lt;/span&gt; nodeselector-pod.yaml
kubectl apply &lt;span class="nt"&gt;-f&lt;/span&gt; affinity-pod.yaml
kubectl get pods &lt;span class="nt"&gt;-o&lt;/span&gt; wide
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the &lt;code&gt;nodeSelector&lt;/code&gt; pod, &lt;code&gt;describe&lt;/code&gt; shows it directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl describe pod nodeselector-pod
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look for a &lt;code&gt;Node-Selectors:&lt;/code&gt; line. It'll show &lt;code&gt;disktype=ssd&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;kubectl describe pod&lt;/code&gt; is useful for inspecting scheduling information and events, but it isn't the best way to inspect the complete affinity configuration. To see the exact affinity rules attached to the Pod, inspect the Pod spec directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl get pod affinity-pod &lt;span class="nt"&gt;-o&lt;/span&gt; yaml | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-A&lt;/span&gt; 16 affinity
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the only reliable way to confirm what's actually attached to a running pod.&lt;/p&gt;

&lt;p&gt;Next up, Part 3: taints and tolerations, where instead of pods choosing nodes, nodes start rejecting pods.&lt;/p&gt;

&lt;p&gt;For those who read till end, I appreciate your patience since this was bit complex, and I too had to struggle to shape this article.  &lt;/p&gt;

&lt;p&gt;If anything here didn't land, or you'd explain it differently, drop a comment. Doubts and pushback are both welcome.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>devops</category>
      <category>beginners</category>
      <category>cloudnative</category>
    </item>
    <item>
      <title>Kubernetes Scheduling 101: What Really Happens Before Your Pod Runs</title>
      <dc:creator>Siddharth More</dc:creator>
      <pubDate>Tue, 15 Sep 2026 05:59:06 +0000</pubDate>
      <link>https://dev.to/siddharthajmore/kubernetes-scheduling-101-what-really-happens-before-your-pod-runs-3khk</link>
      <guid>https://dev.to/siddharthajmore/kubernetes-scheduling-101-what-really-happens-before-your-pod-runs-3khk</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Scheduling means binding a pod to a node by setting &lt;code&gt;spec.nodeName&lt;/code&gt;. You can set this yourself (manual scheduling), or let kube-scheduler do it.&lt;/li&gt;
&lt;li&gt;kube-scheduler decides in two steps: filtering out nodes that can't run the pod, then scoring the ones that can.&lt;/li&gt;
&lt;li&gt;A pod stays on its node for life once scheduled, with one exception: a &lt;code&gt;NoExecute&lt;/code&gt; taint can evict it later. More on that in Part 3.&lt;/li&gt;
&lt;li&gt;Manual scheduling via &lt;code&gt;nodeName&lt;/code&gt; skips the filter-and-score process entirely, so it also skips the safety checks that come with it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You run &lt;code&gt;kubectl apply&lt;/code&gt;, and the pod just shows up somewhere. But how does Kubernetes actually pick that node?&lt;/p&gt;

&lt;p&gt;This is Part 1 of a series on Kubernetes scheduling. Before we get into node selectors, taints, and affinity rules, we need to understand what scheduling actually is. And the easiest way to understand it is to see what happens when you skip it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What scheduling actually means
&lt;/h2&gt;

&lt;p&gt;Scheduling comes down to one thing: binding a pod to a node. That means writing the node's name into the pod's &lt;code&gt;spec.nodeName&lt;/code&gt; field. Once that field is set, the kubelet on that node picks up the pod and runs it.&lt;/p&gt;

&lt;p&gt;Here's the part most people miss. This binding happens once, at creation time. Kubernetes doesn't sit there watching your pods and moving them around as conditions change.&lt;/p&gt;

&lt;p&gt;Once a pod lands on a node, it stays there for its lifetime. Even if a better node shows up five minutes later.&lt;/p&gt;

&lt;p&gt;* There's one exception to that "stays there forever" claim. Taints with the &lt;code&gt;NoExecute&lt;/code&gt; effect can actively kick a running pod off a node if it doesn't tolerate that taint. We'll get into &lt;code&gt;NoExecute&lt;/code&gt; in Part 3. For now, just keep that asterisk in the back of your mind.&lt;/p&gt;

&lt;h2&gt;
  
  
  Manual scheduling: setting &lt;code&gt;nodeName&lt;/code&gt; yourself
&lt;/h2&gt;

&lt;p&gt;Since binding a pod to a node is just setting a field, you can set it yourself.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Pod&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;manually-placed-pod&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;nodeName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;k8s-worker-1&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Apply that, and the pod goes straight to &lt;code&gt;k8s-worker-1&lt;/code&gt;. No decision-making involved. This is literally what the docs call manual scheduling.&lt;/p&gt;

&lt;p&gt;But here's the catch. When you set &lt;code&gt;nodeName&lt;/code&gt; yourself, you skip the whole decision-making process. And that costs you a few safety checks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No resource check. If &lt;code&gt;k8s-worker-1&lt;/code&gt; doesn't have enough CPU or memory, the pod doesn't get rejected. It just fails to run properly, or gets stuck.&lt;/li&gt;
&lt;li&gt;No existence check. Typo the node name, and there's no scheduler around to catch it. The pod sits in &lt;code&gt;Pending&lt;/code&gt;, and the reason isn't obvious at a glance.&lt;/li&gt;
&lt;li&gt;No taint awareness. If &lt;code&gt;k8s-worker-1&lt;/code&gt; has a taint the pod doesn't tolerate, manual placement ignores it completely.
It's a handy tool for demos and edge cases. But it's not something you build on.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;📖 Docs: &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/#nodename" rel="noopener noreferrer"&gt;Assigning pods to nodes using &lt;strong&gt;&lt;em&gt;nodeName&lt;/em&gt;&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Enter kube-scheduler: the two-phase process
&lt;/h2&gt;

&lt;p&gt;For every other pod, &lt;code&gt;nodeName&lt;/code&gt; starts out empty. Filling it in is kube-scheduler's job. It does that in two steps.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Filtering.&lt;/strong&gt; kube-scheduler looks at every node and throws out the ones that can't run the pod at all. Not enough CPU or memory? Out. A taint the pod doesn't tolerate? Out. Port conflict, volume mismatch, whatever hard requirement it fails? Out.&lt;/p&gt;

&lt;p&gt;Whatever's left after that is the feasible set. Every node in it can technically run the pod.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scoring.&lt;/strong&gt; Now kube-scheduler ranks those feasible nodes against each other. It might favor nodes with more free resources, to spread load around. It might favor a node that already has the image cached, so the pod starts faster. There are other scoring rules too, depending on how your cluster's set up.&lt;/p&gt;

&lt;p&gt;Whichever node scores highest wins. That's when &lt;code&gt;nodeName&lt;/code&gt; actually gets set - by the scheduler, not by you.&lt;/p&gt;

&lt;p&gt;This filter-then-score process is exactly what manual scheduling skips. Node selectors, taints, affinity, all of it works by shaping what goes into this process. None of it replaces it.&lt;/p&gt;

&lt;p&gt;📖 Docs: &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/kube-scheduler/#kube-scheduler-implementation" rel="noopener noreferrer"&gt;kube-scheduler's Filter-then-Score process&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Demo: manual vs scheduler-driven placement
&lt;/h2&gt;

&lt;p&gt;Let's see this side by side. Deploy two pods that are basically identical, except one has &lt;code&gt;nodeName&lt;/code&gt; set and one doesn't.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# manual-pod.yaml&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Pod&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;manual-pod&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;nodeName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;k8s-worker-1&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="c1"&gt;# scheduled-pod.yaml&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Pod&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;scheduled-pod&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Apply both. Then check where they landed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl get pods &lt;span class="nt"&gt;-o&lt;/span&gt; wide
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both should be &lt;code&gt;Running&lt;/code&gt;, maybe on different nodes. Now look at the events.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl describe pod scheduled-pod
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You'll see a &lt;code&gt;Scheduled&lt;/code&gt; event, something like &lt;code&gt;Successfully assigned default/scheduled-pod to &amp;lt;node&amp;gt;&lt;/code&gt;. That's a record of the scheduler actually making a decision.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl describe pod manual-pod
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No &lt;code&gt;Scheduled&lt;/code&gt; event at all. As far as the scheduler's concerned, it was never involved. The kubelet on &lt;code&gt;k8s-worker-1&lt;/code&gt; just picked up a pod that already had its name written on it.&lt;/p&gt;

&lt;p&gt;That missing event is the clearest proof you'll get that these are two different code paths, not two flavors of the same thing.&lt;/p&gt;

&lt;p&gt;If you don't see it there, events expire after a while and &lt;code&gt;describe&lt;/code&gt; won't always show old ones. Run &lt;code&gt;kubectl get events&lt;/code&gt; instead and look for the &lt;code&gt;Scheduled&lt;/code&gt; reason.&lt;/p&gt;

&lt;h2&gt;
  
  
  Manual scheduling, and what it skips
&lt;/h2&gt;

&lt;p&gt;Calling &lt;code&gt;nodeName&lt;/code&gt; "manual scheduling" is fine, and it's the term you'll see everywhere, including in CKA material. The label isn't the important part.&lt;/p&gt;

&lt;p&gt;What matters is what it skips. &lt;code&gt;nodeName&lt;/code&gt; doesn't run through kube-scheduler's filter-then-score engine at all, it just hands the kubelet a node and says go.&lt;/p&gt;

&lt;p&gt;That's the distinction that matters for the rest of this series. Node selectors, node affinity, taints and tolerations, pod affinity, none of them bypass the scheduler. They all work with it, either by narrowing the feasible set or nudging the score. Keep that in mind and the rest of this series will click faster.&lt;/p&gt;

&lt;p&gt;Next up, Part 2: node selectors and node affinity, the simplest way to actually influence the scheduler instead of skipping it.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>beginners</category>
      <category>devops</category>
    </item>
    <item>
      <title>Kubernetes and Swap: Why It's Disabled by Default (And How That's Changing) - A kubeadm Debugging Story</title>
      <dc:creator>Siddharth More</dc:creator>
      <pubDate>Mon, 07 Sep 2026 17:32:53 +0000</pubDate>
      <link>https://dev.to/siddharthajmore/kubernetes-and-swap-why-its-disabled-by-default-and-how-thats-changing-a-kubeadm-debugging-g26</link>
      <guid>https://dev.to/siddharthajmore/kubernetes-and-swap-why-its-disabled-by-default-and-how-thats-changing-a-kubeadm-debugging-g26</guid>
      <description>&lt;p&gt;If you've ever bootstrapped a Kubernetes cluster from scratch, you've probably hit this error at least once:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;err="failed to run Kubelet: running with swap on is not supported, please disable swap or set --fail-swap-on flag to false"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I ran into it while setting up my own practice cluster on &lt;strong&gt;LXC containers&lt;/strong&gt; using &lt;strong&gt;kubeadm&lt;/strong&gt;. Everything else checked out - containerd (container runtime), cgroup drivers, network config, but kubelet flatly refused to start. The reason: swap was still enabled.&lt;/p&gt;

&lt;p&gt;This post walks through &lt;em&gt;why&lt;/em&gt; that happens, an LXC-specific nuance worth knowing, and what's actually changed in recent Kubernetes releases around swap support.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick Refresher: What Is Swap?
&lt;/h2&gt;

&lt;p&gt;Swap is disk space (a partition or file) that the Linux kernel uses as overflow when physical RAM runs out, moving inactive memory pages to disk to free up RAM - slower than RAM, but it acts as a safety net against out-of-memory crashes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Kubernetes Disables Swap By Default
&lt;/h2&gt;

&lt;p&gt;By default, &lt;code&gt;kubelet&lt;/code&gt; requires swap to be turned off on every node. If swap is on, kubelet won't even start, and you'll get exactly the error above.&lt;/p&gt;

&lt;p&gt;The reasoning is about predictability. Kubernetes' scheduler and eviction system rely on accurate memory accounting to decide where to place pods and when to evict them under pressure. Swap breaks that: a pod could look like it's within its memory limit while actually thrashing against disk-backed swap, invisible to the scheduler. So historically, the safest default was simple: no swap, period.&lt;/p&gt;

&lt;h2&gt;
  
  
  The LXC-Specific Gotcha: Swap Isn't Namespaced
&lt;/h2&gt;

&lt;p&gt;Here's a detail worth knowing if you're running Kubernetes nodes as &lt;strong&gt;LXC containers&lt;/strong&gt; rather than full VMs: swap is a &lt;strong&gt;global, host-wide kernel resource&lt;/strong&gt; - not namespaced like CPU, memory limits, or networking.&lt;/p&gt;

&lt;p&gt;So &lt;code&gt;/proc/swaps&lt;/code&gt; inside an LXC container reflects the &lt;strong&gt;host's&lt;/strong&gt; swap state, not a per-container one. If the host has swap enabled, every LXC container on it - including your Kubernetes node - sees swap as "on" regardless of container-level settings. You can't hide it from the container alone.&lt;/p&gt;

&lt;p&gt;Your real options in this setup:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Disable swap on the host&lt;/strong&gt; (&lt;code&gt;swapoff -a&lt;/code&gt; + comment out the entry in &lt;code&gt;/etc/fstab&lt;/code&gt;) - the cleanest fix. (&lt;em&gt;But I won't recommend&lt;/em&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tell kubelet to tolerate swap being present&lt;/strong&gt; via config (see next section) (&lt;em&gt;Recommended&lt;/em&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Restrict swap at the cgroup level&lt;/strong&gt; (&lt;code&gt;memory.swap.max=0&lt;/code&gt; on cgroup v2) as an extra safety net - though it limits &lt;em&gt;usage&lt;/em&gt;, not detection, so kubelet still needs to be told to tolerate it.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  How I Fixed It
&lt;/h2&gt;

&lt;p&gt;Since I wanted to move fast without reprovisioning a separate swap-free host, I set &lt;code&gt;failSwapOn: false&lt;/code&gt; in the kubelet configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# /var/lib/kubelet/config.yaml&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kubelet.config.k8s.io/v1beta1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;KubeletConfiguration&lt;/span&gt;
&lt;span class="na"&gt;failSwapOn&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Restart kubelet, and it starts up right away - running in &lt;code&gt;NoSwap&lt;/code&gt; mode by default.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bigger Picture: KEP-2400 and Swap Support Going GA
&lt;/h2&gt;

&lt;p&gt;While debugging this, I learned that Kubernetes' stance on swap has actually been evolving for years - and it's more nuanced than "swap = broken."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;a href="https://github.com/kubernetes/enhancements/tree/master/keps/sig-node/2400-node-swap" rel="noopener noreferrer"&gt;KEP-2400 (Node Memory Swap Support)&lt;/a&gt;&lt;/strong&gt; started as alpha in Kubernetes 1.22 (2021) and reached &lt;strong&gt;GA in Kubernetes 1.34&lt;/strong&gt; (2025). It introduces the &lt;code&gt;NodeSwap&lt;/code&gt; feature, controlled via kubelet configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;failSwapOn&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;span class="na"&gt;memorySwap&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;swapBehavior&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;LimitedSwap&lt;/span&gt;   &lt;span class="c1"&gt;# or NoSwap&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The two supported behaviors:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;NoSwap&lt;/code&gt;&lt;/strong&gt; - the default. Kubernetes workloads (pods) don't and can't use swap, but processes outside Kubernetes' scope - system daemons, and even kubelet itself - still can. This protects the node from system-level memory spikes without giving that safety net to workloads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;LimitedSwap&lt;/code&gt;&lt;/strong&gt; - pods get a bounded, proportional amount of swap based on their memory &lt;em&gt;requests&lt;/em&gt;, via cgroup v2. Guaranteed and BestEffort QoS pods get zero swap by design; only Burstable pods (requests &amp;lt; limits) get meaningful access.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An earlier mode, &lt;code&gt;UnlimitedSwap&lt;/code&gt;, was removed in 1.30 - it let a single misbehaving container consume unlimited swap and potentially crash the node.&lt;/p&gt;

&lt;h3&gt;
  
  
  GA Doesn't Mean "Auto-Enabled"
&lt;/h3&gt;

&lt;p&gt;Going GA doesn't mean kubelet now silently tolerates swap on upgrade. On the latest version, swap left on with no config changes still triggers the same startup error - the safe default hasn't moved. What's changed is there's now a &lt;strong&gt;stable, supported way to opt in&lt;/strong&gt;, useful for things like JVM-based apps (Jenkins, SonarQube) that reserve large amounts of memory but rarely touch all of it, or as a buffer against short memory spikes instead of an outright OOMKill.&lt;/p&gt;

&lt;p&gt;Swap on Kubernetes nodes is still an advanced, deliberate choice - control-plane nodes in particular are recommended to stay swap-free.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Kubelet requires swap off by default - unchanged even with KEP-2400 reaching GA in 1.34.&lt;/li&gt;
&lt;li&gt;On LXC containers, remember swap is a host-level kernel resource, not namespaced.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;failSwapOn: false&lt;/code&gt; gets kubelet running quickly (default &lt;code&gt;NoSwap&lt;/code&gt;), but disabling swap at the host level is still the standard fix for a &lt;strong&gt;dedicated&lt;/strong&gt; node.&lt;/li&gt;
&lt;li&gt;If you want to use swap intentionally, &lt;code&gt;LimitedSwap&lt;/code&gt; gives a bounded, QoS-aware way to do it - worth exploring for memory-heavy, low-active-usage workloads.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is my first dev.to post, so if I got something wrong or oversimplified, or if you've handled swap differently on your clusters, I'd genuinely appreciate the correction - drop a comment below.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>linux</category>
      <category>devops</category>
      <category>kubeadm</category>
    </item>
  </channel>
</rss>
