<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Roy</title>
    <description>The latest articles on DEV Community by Roy (@roylib).</description>
    <link>https://dev.to/roylib</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4061547%2Fe7832376-d1a0-48b0-a141-6c36e523c022.png</url>
      <title>DEV Community: Roy</title>
      <link>https://dev.to/roylib</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/roylib"/>
    <language>en</language>
    <item>
      <title>Karpenter Troubleshooting: Pending Pods, Failed Launches, and Nodes That Never Join</title>
      <dc:creator>Roy</dc:creator>
      <pubDate>Mon, 28 Sep 2026 17:53:11 +0000</pubDate>
      <link>https://dev.to/skyhook-radar/karpenter-troubleshooting-pending-pods-failed-launches-and-nodes-that-never-join-467h</link>
      <guid>https://dev.to/skyhook-radar/karpenter-troubleshooting-pending-pods-failed-launches-and-nodes-that-never-join-467h</guid>
      <description>&lt;p&gt;&lt;a href="https://karpenter.sh/" rel="noopener noreferrer"&gt;Karpenter&lt;/a&gt; changes how you debug a Pending pod.&lt;/p&gt;

&lt;p&gt;A Karpenter NodePool can report &lt;code&gt;Ready=True&lt;/code&gt; while the nodes it launches repeatedly fail to register. And that is only one reason these incidents can be confusing: the scheduler, NodePool, NodeClaim, NodeClass, and controller logs each hold a different part of the answer.&lt;/p&gt;

&lt;p&gt;On a fixed set of nodes, Pending usually means the nodes you already have cannot take the pod. With Karpenter, there is another question: did Karpenter rule out every NodePool, or did it try to create capacity and fail?&lt;/p&gt;

&lt;p&gt;This guide maps the failures we run into most often to the place I would look next. The screenshots come from Radar's Capacity view, the open-source tool we maintain, which assembles the same evidence from the cluster itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Karpenter pending pods: symptom, cause, where to look
&lt;/h2&gt;

&lt;p&gt;Scheduler messages usually arrive as &lt;code&gt;0/N nodes are available:&lt;/code&gt; followed by a reason. That message is useful, but with Karpenter it is often only the first step.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symptom or message&lt;/th&gt;
&lt;th&gt;Usual cause&lt;/th&gt;
&lt;th&gt;Where to look&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;didn't match Pod's node affinity/selector&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Pod requires a label no NodePool declares&lt;/td&gt;
&lt;td&gt;NodePool &lt;code&gt;requirements&lt;/code&gt; and &lt;code&gt;spec.template.metadata.labels&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;had untolerated taint {key: value}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;NodePool template taints the node, pod has no toleration&lt;/td&gt;
&lt;td&gt;NodePool &lt;code&gt;spec.template.spec.taints&lt;/code&gt; vs pod tolerations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;Insufficient cpu&lt;/code&gt; / &lt;code&gt;Insufficient memory&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Existing nodes are full. Says nothing about what Karpenter could provision&lt;/td&gt;
&lt;td&gt;Node allocatable vs scheduled requests, then Karpenter's scheduling event&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pod pending, no NodeClaim created&lt;/td&gt;
&lt;td&gt;Every pool ruled the pod out, or NodePool limits are reached&lt;/td&gt;
&lt;td&gt;Karpenter &lt;code&gt;FailedScheduling&lt;/code&gt; event; &lt;code&gt;status.resources&lt;/code&gt; vs &lt;code&gt;spec.limits&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NodeClaim created, no node appears&lt;/td&gt;
&lt;td&gt;Launch failed, or instance launched and never registered&lt;/td&gt;
&lt;td&gt;NodeClaim conditions for launch failures; &lt;code&gt;NodeRegistrationHealthy&lt;/code&gt; for repeated registration failures&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nodes launch then disappear&lt;/td&gt;
&lt;td&gt;Registration timeout&lt;/td&gt;
&lt;td&gt;NodePool &lt;code&gt;NodeRegistrationHealthy&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pod stuck with no scheduler attempt&lt;/td&gt;
&lt;td&gt;Scheduling gate still present&lt;/td&gt;
&lt;td&gt;&lt;code&gt;spec.schedulingGates&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nodes not consolidating&lt;/td&gt;
&lt;td&gt;Disruption blocked or budgets exhausted&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;DisruptionBlocked&lt;/code&gt; / &lt;code&gt;Unconsolidatable&lt;/code&gt; events&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9se5mczwhh7b8k3f9lop.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9se5mczwhh7b8k3f9lop.png" alt="The Capacity overview: capacity managers, the cluster scheduling bar, prioritized operational signals including " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The important part is that these failures do not all live on the same object. A scheduling event can tell you why a pool was rejected, while a NodeClaim or NodePool condition tells you why capacity that should have appeared never did.&lt;/p&gt;

&lt;h2&gt;
  
  
  "didn't match Pod's node affinity/selector"
&lt;/h2&gt;

&lt;p&gt;This is a common Karpenter misconfiguration and one of the easier ones to identify from configuration alone.&lt;/p&gt;

&lt;p&gt;Karpenter only applies labels that come from a NodePool's &lt;code&gt;requirements&lt;/code&gt; or its template labels. If a pod requires a custom label that no NodePool declares, no node Karpenter builds will ever carry it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F11wjmno3i9q0w3ofv9n3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F11wjmno3i9q0w3ofv9n3.png" alt="Radar's Demand view showing a blocked group headlined " width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In production this is often a stale label: a &lt;code&gt;workload-type&lt;/code&gt; renamed during a migration, or a &lt;code&gt;nodeSelector&lt;/code&gt; copied from another cluster.&lt;/p&gt;

&lt;p&gt;Start with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl describe pod &amp;lt;pod&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then compare the pod's selectors and required affinity against each NodePool's:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;spec.requirements
spec.template.metadata.labels
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the label the pod requires is never declared by any NodePool, Karpenter has nowhere to place it.&lt;/p&gt;

&lt;p&gt;Radar's Demand view does this comparison across all NodePools and groups pods with the same scheduling constraints together, but the underlying check is the same one you can do manually.&lt;/p&gt;

&lt;h2&gt;
  
  
  "had untolerated taint" on a Karpenter cluster
&lt;/h2&gt;

&lt;p&gt;Karpenter taints nodes from &lt;code&gt;spec.template.spec.taints&lt;/code&gt; on the NodePool.&lt;/p&gt;

&lt;p&gt;A pod without a matching toleration cannot land on anything that pool creates, so Karpenter can rule the pool out before provisioning even starts.&lt;/p&gt;

&lt;p&gt;Compare:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;NodePool.spec.template.spec.taints
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;against the pod's:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;spec.tolerations
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One thing to watch: startup taints are different from permanent taints.&lt;/p&gt;

&lt;p&gt;Startup taints, along with the well-known unreachable/not-ready taints, are expected during parts of a node's lifecycle. They should not automatically be treated as evidence that the NodePool can never run the workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  LaunchFailed, VCPULimitExceeded, and InsufficientInstanceCapacity
&lt;/h2&gt;

&lt;p&gt;If Karpenter creates a NodeClaim and the cloud provider refuses it, the useful evidence moves away from the pod.&lt;/p&gt;

&lt;p&gt;Look at the NodeClaim lifecycle conditions.&lt;/p&gt;

&lt;p&gt;Common AWS examples include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;VCPULimitExceeded&lt;/code&gt; - the account's vCPU quota for that instance family is exhausted.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;InsufficientInstanceCapacity&lt;/code&gt; - AWS does not currently have capacity for that instance type in the requested region or zone.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;Unauthorized&lt;/code&gt; - the instance profile or IAM configuration is wrong.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;LaunchFailed&lt;/code&gt; - Karpenter reached the launch stage but could not create the instance successfully.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A detail that makes these annoying to debug: failed NodeClaims do not necessarily stay around forever.&lt;/p&gt;

&lt;p&gt;Timed-out NodeClaims are deleted, and the conditions attached to them disappear too. Karpenter can surface repeated launch or registration trouble at the NodePool level through &lt;code&gt;NodeRegistrationHealthy&lt;/code&gt;, but the individual history - which claims failed, how many, and in what order - can be harder to reconstruct afterward.&lt;/p&gt;

&lt;p&gt;That is why events and controller logs matter for intermittent failures.&lt;/p&gt;

&lt;p&gt;Radar's Activity view keeps a longer provisioning history, but even without Radar, the practical rule is simple: if you see a NodeClaim, inspect it before assuming the pod event contains the answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why aren't my Karpenter nodes joining the cluster?
&lt;/h2&gt;

&lt;p&gt;This is one of the more confusing failure modes because the NodePool itself can still report &lt;code&gt;Ready=True&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Karpenter tracks &lt;code&gt;NodeRegistrationHealthy&lt;/code&gt; separately from the NodePool's top-level &lt;code&gt;Ready&lt;/code&gt; condition.&lt;/p&gt;

&lt;p&gt;That means the NodePool can be valid for provisioning, successfully launch instances, and still fail to produce usable nodes because those instances never register with the cluster.&lt;/p&gt;

&lt;p&gt;The pattern looks roughly like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Karpenter decides the NodePool can satisfy the workload.&lt;/li&gt;
&lt;li&gt;A NodeClaim is created.&lt;/li&gt;
&lt;li&gt;The cloud instance launches.&lt;/li&gt;
&lt;li&gt;The node never registers successfully.&lt;/li&gt;
&lt;li&gt;The registration timeout is reached.&lt;/li&gt;
&lt;li&gt;The instance is deleted.&lt;/li&gt;
&lt;li&gt;Karpenter tries again.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;From the workload side, the result is simple: the pod stays Pending and capacity never arrives.&lt;/p&gt;

&lt;p&gt;From the NodePool side, &lt;code&gt;Ready=True&lt;/code&gt; alone does not tell the whole story.&lt;/p&gt;

&lt;p&gt;This is different from a broken NodeClass.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fget49qynfhwuu2ekfnvc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fget49qynfhwuu2ekfnvc.png" alt="The gpu-workloads NodePool detail page leading with two warnings - NodeClass Not Ready and NodePool Node Registration Unhealthy - each with a cause and a next step, above the capacity ledger" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For example, if an EC2NodeClass looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ValidationSucceeded=False
SecurityGroupsReady=False
SubnetsReady=False
InstanceProfileReady=Unknown
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;then the NodeClass itself is not ready. That propagates through &lt;code&gt;NodeClassReady&lt;/code&gt;, so this is not a case where the NodePool remains happily &lt;code&gt;Ready&lt;/code&gt; while nodes fail to register.&lt;/p&gt;

&lt;p&gt;Typical NodeClass problems include selectors that resolve to no subnets or security groups, a missing instance profile, or another prerequisite that cannot be resolved before launch.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;Ready=True&lt;/code&gt; plus &lt;code&gt;NodeRegistrationHealthy=False&lt;/code&gt; problem happens later: the NodeClass was valid enough to provision, but the resulting node never successfully joined the cluster.&lt;/p&gt;

&lt;p&gt;Those two failures can look nearly identical from the pod's point of view, but they require looking in different places.&lt;/p&gt;

&lt;h2&gt;
  
  
  Am I about to hit a NodePool limit?
&lt;/h2&gt;

&lt;p&gt;Karpenter stops provisioning when a NodePool reaches &lt;code&gt;spec.limits&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The pods that would have triggered more capacity stay Pending, and no NodeClaim is created for them.&lt;/p&gt;

&lt;p&gt;Karpenter does tell you when this happens. Look for a &lt;code&gt;FailedScheduling&lt;/code&gt; event containing messages such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;node limits have been exhausted for nodepool
all available instance types exceed limits for nodepool
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The number Karpenter checks is &lt;code&gt;status.resources&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That matters because &lt;code&gt;status.resources&lt;/code&gt; includes resources represented by in-flight NodeClaims. So:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;headroom = limit - provisioned
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;headroom = limit - allocatable
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A NodePool can therefore be close to or at its limit even when relatively little capacity has successfully registered with Kubernetes.&lt;/p&gt;

&lt;p&gt;The event tells you which pool hit its limit. To understand how much room is left elsewhere, compare &lt;code&gt;spec.limits&lt;/code&gt; and &lt;code&gt;status.resources&lt;/code&gt; across the other pools.&lt;/p&gt;

&lt;h2&gt;
  
  
  Provisioned, allocatable, requests, usage: which number means what
&lt;/h2&gt;

&lt;p&gt;These numbers answer different questions:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Number&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;What it tells you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Provisioned&lt;/td&gt;
&lt;td&gt;&lt;code&gt;status.resources&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;What counts against NodePool &lt;code&gt;spec.limits&lt;/code&gt;, including in-flight claims&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Node allocatable&lt;/td&gt;
&lt;td&gt;Registered nodes&lt;/td&gt;
&lt;td&gt;What Kubernetes can actually schedule onto now&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scheduled requests&lt;/td&gt;
&lt;td&gt;Pod specs&lt;/td&gt;
&lt;td&gt;What the scheduler has already committed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Actual usage&lt;/td&gt;
&lt;td&gt;Metrics API&lt;/td&gt;
&lt;td&gt;What workloads are consuming, not what can still be scheduled&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Karpenter and Kubernetes schedule based on requests, not current CPU or memory usage.&lt;/p&gt;

&lt;p&gt;A node can be using 20% CPU and still be completely full from the scheduler's point of view if all of its CPU has already been requested.&lt;/p&gt;

&lt;p&gt;That is why Radar keeps eight capacity values separate instead of reducing them to a single utilization percentage.&lt;/p&gt;

&lt;p&gt;There is another subtlety: &lt;code&gt;unallocated&lt;/code&gt; is only an aggregate subtraction.&lt;/p&gt;

&lt;p&gt;If a cluster has 4 CPUs and 8 GiB of memory unallocated in total, that does not prove a pod requesting 4 CPUs and 8 GiB can fit. Those resources still need to exist together on a single node that also satisfies the pod's other constraints.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does DisruptionBlocked mean in Karpenter?
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;DisruptionBlocked&lt;/code&gt; and &lt;code&gt;Unconsolidatable&lt;/code&gt; do not mean Karpenter disrupted a node.&lt;/p&gt;

&lt;p&gt;They mean disruption was &lt;strong&gt;prevented&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Read either as disruption happening and the timeline reads backwards.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftrds6zliq8thrc03lyu5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftrds6zliq8thrc03lyu5.png" alt="An Activity episode reading " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Common causes include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;a disruption budget that is exhausted&lt;/li&gt;
&lt;li&gt;a PodDisruptionBudget that cannot be satisfied&lt;/li&gt;
&lt;li&gt;a &lt;code&gt;karpenter.sh/do-not-disrupt&lt;/code&gt; annotation&lt;/li&gt;
&lt;li&gt;a NodeClaim with no associated node to disrupt&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The raw event is usually the best place to start:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DisruptionBlocked: Nodeclaim does not have an associated node
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not treat these events as evidence that Karpenter removed capacity. They are evidence that Karpenter considered a disruption and could not proceed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a pod evaluates as "unknown" instead of incompatible
&lt;/h2&gt;

&lt;p&gt;When we built this evaluation into Radar, we found it useful to separate three cases:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;declared compatible&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;incompatible&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;unknown&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The difference matters because configuration can prove some scheduling failures, but not all of them.&lt;/p&gt;

&lt;p&gt;A required custom label that no NodePool declares is &lt;strong&gt;incompatible&lt;/strong&gt;. Nothing the cloud provider does can fix that.&lt;/p&gt;

&lt;p&gt;Provider-controlled labels such as zone, instance type, capacity type, and architecture are different. If the NodePool leaves them open, the provider's actual offerings may still satisfy the pod.&lt;/p&gt;

&lt;p&gt;Requests are similar. Observed nodes can tell you whether a known machine shape is large enough, but they do not prove which instance types the provider could create next.&lt;/p&gt;

&lt;p&gt;Radar therefore does not simulate provider offerings or claim that a compatible pod will schedule. It only says that the declared constraints do not rule it out.&lt;/p&gt;

&lt;p&gt;If the available information cannot prove either outcome, the result is &lt;code&gt;unknown&lt;/code&gt; instead of a guess.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why some numbers show ≥ or ?
&lt;/h2&gt;

&lt;p&gt;Capacity information is not always complete.&lt;/p&gt;

&lt;p&gt;For example, RBAC may prevent access to one source, or metrics may only be available for part of the fleet.&lt;/p&gt;

&lt;p&gt;Radar marks that distinction explicitly:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Glyph&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;(none)&lt;/td&gt;
&lt;td&gt;Exact - the source was fully observed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;≥&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Lower bound - only part of the source was observed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;≤&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Upper bound - derived from a lower-bound input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;?&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Unknown - the source was not observed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Missing data is never shown as zero, and partial data is never shown as exact.&lt;/p&gt;

&lt;p&gt;This matters during an incident because "0" and "I could not read it" are very different answers.&lt;/p&gt;

&lt;h2&gt;
  
  
  What this doesn't cover
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No full scheduling simulation.&lt;/strong&gt; Provider inventory, spot availability, and exact bin-packing remain Karpenter's job.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DRA demand is outside the normal requests ledger.&lt;/strong&gt; Dynamic Resource Allocation &lt;code&gt;ResourceClaims&lt;/code&gt; can carry accelerator demand outside normal container requests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No long-term trend analysis.&lt;/strong&gt; This is about current state and recent failure evidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Single cluster.&lt;/strong&gt; Same scope as the rest of Radar OSS.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Clusters without Karpenter can still be inspected through the normal node and node-group views. The Karpenter-specific Demand and Activity analysis depends on NodePools and NodeClaims being present.&lt;/p&gt;

&lt;h2&gt;
  
  
  Running it
&lt;/h2&gt;

&lt;p&gt;Radar is Apache-2.0, runs as a single binary, and does not install anything in the cluster:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;brew &lt;span class="nb"&gt;install &lt;/span&gt;skyhook-io/tap/radar
radar
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Capacity requires Radar v1.9.0 or newer.&lt;/p&gt;

&lt;p&gt;You can open it from the sidebar or jump there directly from a stuck pod using "Evaluate against Karpenter NodePools".&lt;/p&gt;

&lt;p&gt;If Radar reaches a verdict you can prove wrong, that's worth an &lt;a href="https://github.com/skyhook-io/radar" rel="noopener noreferrer"&gt;issue on GitHub&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Related reading: &lt;a href="https://radarhq.io/blog/five-questions-you-cant-answer-with-kubectl" rel="noopener noreferrer"&gt;Five Questions You Can't Answer With kubectl&lt;/a&gt;, &lt;a href="https://radarhq.io/blog/kubernetes-network-reachability" rel="noopener noreferrer"&gt;Everything Is Green and Nothing Works: Network Reachability in Radar&lt;/a&gt;, and &lt;a href="https://radarhq.io/blog/kubernetes-1-37-breaking-changes" rel="noopener noreferrer"&gt;Kubernetes 1.37 Breaking Changes&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>karpenter</category>
      <category>devops</category>
      <category>sre</category>
    </item>
  </channel>
</rss>
