<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: yosshi_</title>
    <description>The latest articles on DEV Community by yosshi_ (@yosshi_).</description>
    <link>https://dev.to/yosshi_</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4167980%2F83ccb051-b6c0-43c0-b5db-451d1e502407.jpg</url>
      <title>DEV Community: yosshi_</title>
      <link>https://dev.to/yosshi_</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/yosshi_"/>
    <language>en</language>
    <item>
      <title>Kubernetes Memory Usage and cgroup v2: Telling OOM Kills from Evictions and Right-Sizing Memory Requests</title>
      <dc:creator>yosshi_</dc:creator>
      <pubDate>Fri, 09 Oct 2026 07:20:35 +0000</pubDate>
      <link>https://dev.to/yosshi_/kubernetes-memory-usage-and-cgroup-v2-telling-oom-kills-from-evictions-and-right-sizing-memory-1bi3</link>
      <guid>https://dev.to/yosshi_/kubernetes-memory-usage-and-cgroup-v2-telling-oom-kills-from-evictions-and-right-sizing-memory-1bi3</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;This is an English translation of &lt;a href="https://qiita.com/yosshi_/items/46d8425b188675c102e5" rel="noopener noreferrer"&gt;an article I originally published in Japanese on Qiita&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Scope: Kubernetes v1.37, Linux 6.12, cgroup v2, containerd. Exact versions are listed in Components and versions.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;"Memory usage" in Kubernetes is several different numbers, and &lt;strong&gt;OOM kills and evictions are each triggered by a different one of them&lt;/strong&gt;. That is why a container can be &lt;code&gt;OOMKilled&lt;/code&gt; while &lt;code&gt;kubectl top&lt;/code&gt; shows headroom, or a Pod can be evicted while node utilization looks low.&lt;/p&gt;

&lt;p&gt;Terms used throughout (details): &lt;strong&gt;memcg OOM&lt;/strong&gt; is an OOM kill for exceeding a configured limit (container, Pod or &lt;code&gt;/kubepods.slice&lt;/code&gt;), and &lt;strong&gt;global OOM&lt;/strong&gt; is an OOM kill caused by memory exhaustion of the whole node. &lt;strong&gt;Eviction&lt;/strong&gt; is a separate mechanism, done by the kubelet.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Event&lt;/th&gt;
&lt;th&gt;Triggering value&lt;/th&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
memcg OOM (OOM kill from exceeding a limit)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;memory.current&lt;/code&gt; (&lt;code&gt;container_memory_usage_bytes&lt;/code&gt;) of the container, Pod or &lt;code&gt;/kubepods.slice&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Reaches &lt;code&gt;memory.max&lt;/code&gt; at the same level &lt;strong&gt;and reclaim cannot bring it back under&lt;/strong&gt; (a container's &lt;code&gt;memory.max&lt;/code&gt; is its memory limit)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
global OOM (OOM kill from node memory exhaustion)&lt;/td&gt;
&lt;td&gt;Whether the whole node's memory is exhausted (from the kernel's point of view)&lt;/td&gt;
&lt;td&gt;There is no threshold. The victim is chosen by &lt;code&gt;oom_score&lt;/code&gt; (QoS class and memory request)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Eviction trigger&lt;/td&gt;
&lt;td&gt;The node's &lt;code&gt;memory.available&lt;/code&gt; (= &lt;code&gt;MemTotal&lt;/code&gt; − node working set)&lt;/td&gt;
&lt;td&gt;Capacity falls below the &lt;code&gt;evictionHard&lt;/code&gt; / &lt;code&gt;evictionSoft&lt;/code&gt; thresholds (kubelet default: &lt;code&gt;memory.available&amp;lt;100Mi&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Choosing which Pod to evict&lt;/td&gt;
&lt;td&gt;Difference between a Pod's working set and its memory request&lt;/td&gt;
&lt;td&gt;Pods whose actual usage exceeds their request are preferred&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Key points:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;kubectl top&lt;/code&gt; and &lt;code&gt;container_memory_working_set_bytes&lt;/code&gt; show the working set, not RSS.&lt;/strong&gt; It includes active page cache, &lt;code&gt;shmem&lt;/code&gt; and kernel memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;memcg OOM compares &lt;code&gt;memory.current&lt;/code&gt; against the limit, not the working set&lt;/strong&gt;, and only kills when reclaim cannot bring usage back under the limit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Eviction is decided from the node-wide working set, not &lt;code&gt;MemAvailable&lt;/code&gt;&lt;/strong&gt;, so node-exporter utilization can look healthy while the kubelet is close to its threshold.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;OOMKilled&lt;/code&gt; does not tell you the path.&lt;/strong&gt; A &lt;code&gt;SystemOOM&lt;/code&gt; event on the Node means global OOM, but its absence proves nothing (events are kept for 1 hour by default).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The most effective fix for global OOM and eviction is to set the memory request to match actual usage.&lt;/strong&gt; It does not help with memcg OOM, which needs a limit review.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Where to go next:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why do the numbers disagree? Quick reference and Kubernetes memory usage and cgroup v2
&lt;/li&gt;
&lt;li&gt;Something was just killed or evicted: Investigating OOM and eviction
&lt;/li&gt;
&lt;li&gt;Prevent it: What to do about it, with PromQL for checking requests&lt;/li&gt;
&lt;li&gt;Specific questions: FAQ
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;I get asked about Kubernetes memory usage again and again, and I kept giving the same explanation. The information you need is spread across the Kubernetes and Linux kernel documentation, and I could not find anything that connects the two, so I wrote it down myself.&lt;/p&gt;
&lt;h3&gt;
  
  
  There is no single definition of "memory usage"
&lt;/h3&gt;

&lt;p&gt;The question "how many MiB of memory is this container using?" &lt;strong&gt;has no single correct answer&lt;/strong&gt;. Unlike CPU utilization or disk usage, memory usage is not one value with one definition. &lt;strong&gt;The number changes depending on what you want to measure and where you measure it from.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The reason is that Linux cannot assign every byte of memory to exactly one owner.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Who owns the page cache?&lt;/strong&gt; When you read a file, the kernel keeps its contents in memory. The application did not request this with &lt;code&gt;malloc&lt;/code&gt;; the kernel allocated it on its own for performance. And when memory runs short, the kernel silently throws it away. Whether it counts as "memory the application is using" depends on your purpose.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Who owns shared memory?&lt;/strong&gt; Shared library pages, tmpfs and shared memory are referenced by several processes through the same physical pages. If you simply add up per-process numbers, the same page is counted many times, and the total can even exceed physical memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;When does "freed" memory actually go down?&lt;/strong&gt; Even if an application calls &lt;code&gt;free&lt;/code&gt;, the OS-visible usage does not drop unless the allocator or language runtime returns the memory to the OS. If it does return it, usage drops. The value depends on which layer you look at.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Some memory is reclaimable and some is not.&lt;/strong&gt; Both count as "in use", but under memory pressure the kernel can take some of it back (page cache) and cannot take back the rest (anonymous pages). You cannot ignore this distinction when estimating free memory.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In other words, memory usage is &lt;strong&gt;a value that only becomes defined once you fix the measurement method and the purpose&lt;/strong&gt;. "The pages a process holds in physical memory", "the memory only that process uses" and "the memory that cannot be freed under pressure" are all legitimate definitions, and each gives a different number.&lt;/p&gt;
&lt;h3&gt;
  
  
  Kubernetes runs on its own definitions
&lt;/h3&gt;

&lt;p&gt;On top of that, &lt;strong&gt;Kubernetes components use their own definitions for management purposes&lt;/strong&gt;. If you read a value assuming it is the same thing you normally call "memory usage", you will misinterpret it.&lt;/p&gt;

&lt;p&gt;The container memory usage you see with &lt;code&gt;kubectl top pod&lt;/code&gt; or &lt;code&gt;container_memory_working_set_bytes&lt;/code&gt; is &lt;strong&gt;not the amount of memory the application has allocated (RSS)&lt;/strong&gt;. Kubernetes treats the working set as container usage, and it includes RSS plus part of the page cache and memory used by the kernel.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;container_memory_working_set_bytes      &lt;span class="c"&gt;# the value kubectl top pod shows&lt;/span&gt;
  ≒ container_memory_rss                &lt;span class="c"&gt;# memory the application allocated (anonymous pages)&lt;/span&gt;
  + active page cache                   &lt;span class="c"&gt;# only inactive pages are subtracted&lt;/span&gt;
  + tmpfs and shared memory &lt;span class="o"&gt;(&lt;/span&gt;shmem&lt;span class="o"&gt;)&lt;/span&gt;     &lt;span class="c"&gt;# included in full, active or inactive&lt;/span&gt;
  + kernel memory, etc.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;So the working set is larger than what the application actually consumes. The tricky part is that &lt;strong&gt;OOM kills and evictions happen because a value other than the one you are looking at reached its limit&lt;/strong&gt;. For an OOM kill caused by exceeding a limit (called memcg OOM in this article), the kernel compares the container's &lt;code&gt;memory.current&lt;/code&gt; against the limit, which is not the working set. And &lt;code&gt;memory.current&lt;/code&gt; merely reaching the limit does not cause an OOM; the kill happens when reclaim fails to bring usage back under the limit. Eviction is triggered by free memory computed from the node-wide working set, so it looks at a different scope than the per-container values.&lt;/p&gt;

&lt;p&gt;If you treat these as the same thing, you cannot explain situations that look contradictory:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The working set has not reached the memory limit, yet the container was &lt;code&gt;OOMKilled&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Node memory utilization computed from &lt;code&gt;node_memory_MemAvailable_bytes&lt;/code&gt; is low, yet a Pod was evicted&lt;/li&gt;
&lt;li&gt;The application does not leak memory, yet the working set keeps growing&lt;/li&gt;
&lt;li&gt;The RSS the application reports does not match &lt;code&gt;container_memory_rss&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these are metric bugs. They are the result of &lt;strong&gt;each component reporting a value with a definition that suits its own purpose&lt;/strong&gt;. Conversely, once you know which value has which definition, the apparent contradictions can be explained.&lt;/p&gt;

&lt;p&gt;This article covers three things in order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Kubernetes memory usage and cgroup v2&lt;/strong&gt;: what each value that appears as "memory usage" measures, which cgroup v2 value it corresponds to, and which value triggers eviction&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;How to investigate OOM kills and evictions&lt;/strong&gt;: there are two OOM kill paths, and how to tell which one happened when you see &lt;code&gt;OOMKilled&lt;/code&gt; or an eviction&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What to do about it&lt;/strong&gt;: what to do once you know the cause. &lt;strong&gt;Right-sizing the memory request is the most effective way to reduce OOM kills and evictions together&lt;/strong&gt;, so the article focuses on that, including PromQL for checking it&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;
  
  
  How this article names the two OOM kill paths
&lt;/h2&gt;

&lt;p&gt;There are two paths by which a Pod's process gets killed for memory reasons, and &lt;strong&gt;they differ in both cause and blast radius&lt;/strong&gt;. If you call both "OOM" you cannot tell them apart, so &lt;strong&gt;this article uses the following two names throughout&lt;/strong&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Name in this article&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;th&gt;Triggering cause&lt;/th&gt;
&lt;th&gt;What gets killed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;memcg OOM&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Enforcing a configured limit&lt;/td&gt;
&lt;td&gt;A container, Pod or &lt;code&gt;/kubepods.slice&lt;/code&gt; &lt;strong&gt;exceeded its &lt;code&gt;memory.max&lt;/code&gt; (= memory limit)&lt;/strong&gt;
&lt;/td&gt;
&lt;td&gt;Only processes inside the cgroup that hit its limit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;global OOM&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Keeping the whole system alive (protecting the node)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Memory exhaustion of the whole node&lt;/strong&gt;, regardless of any container's limit&lt;/td&gt;
&lt;td&gt;Any process on the node&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Put simply, &lt;strong&gt;memcg OOM is caused by "your own limit", and global OOM is caused by "the node's free memory"&lt;/strong&gt;. Many cases of &lt;code&gt;OOMKilled&lt;/code&gt; despite headroom under the limit are the latter.&lt;/p&gt;

&lt;p&gt;The difference in purpose is written in the kernel documentation. Global OOM exists &lt;strong&gt;to keep the rest of the system alive&lt;/strong&gt;; &lt;a href="https://docs.kernel.org/admin-guide/mm/concepts.html#oom-killer" rel="noopener noreferrer"&gt;Concepts overview - OOM killer&lt;/a&gt; says "In order to save the rest of the system, it invokes the &lt;code&gt;OOM killer&lt;/code&gt;". The description of &lt;code&gt;memory.max&lt;/code&gt;, on the other hand, is "This is the main mechanism to limit memory usage of a cgroup" (&lt;a href="https://docs.kernel.org/admin-guide/cgroup-v2.html#memory-interface-files" rel="noopener noreferrer"&gt;Memory Interface Files&lt;/a&gt;), with no mention of protecting other cgroups. cgroup v2 treats limits (&lt;code&gt;memory.max&lt;/code&gt; / &lt;code&gt;memory.high&lt;/code&gt;) and protection (&lt;code&gt;memory.min&lt;/code&gt; / &lt;code&gt;memory.low&lt;/code&gt;) as separate models ("implements both limit and protection models"), and memcg OOM belongs to the former.&lt;/p&gt;

&lt;p&gt;So &lt;strong&gt;memcg OOM can happen even when the node has plenty of free memory&lt;/strong&gt;. It happens solely because a configured limit was exceeded, regardless of actual pressure. Global OOM, conversely, happens only when node memory is genuinely exhausted, regardless of any limit.&lt;/p&gt;

&lt;p&gt;The one exception is the &lt;code&gt;/kubepods.slice&lt;/code&gt; hierarchy. Its &lt;code&gt;memory.max&lt;/code&gt; is set to &lt;code&gt;Capacity - kube-reserved - system-reserved&lt;/code&gt; (see below), which effectively reserves memory for non-Pod processes such as the kubelet and containerd. That is not the purpose of the kernel's memcg OOM; it is Kubernetes using the mechanism that way.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note: These two names are not official Kubernetes terms.&lt;/strong&gt; Kubernetes only exposes the termination reason &lt;code&gt;OOMKilled&lt;/code&gt;, and it does not distinguish the paths (the reason is explained in Why OOMKilled cannot distinguish the paths).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Both are terms the kernel itself uses, and this article adopts them as they are.&lt;/strong&gt; The two are used side by side in a kernel comment (&lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/oom_kill.c#L1178-L1183" rel="noopener noreferrer"&gt;mm/oom_kill.c&lt;/a&gt;, &lt;code&gt;pagefault_out_of_memory()&lt;/code&gt;):&lt;/p&gt;


&lt;pre class="highlight c"&gt;&lt;code&gt;&lt;span class="cm"&gt;/*
 * The pagefault handler calls here because some allocation has failed. We have
 * to take care of the memcg OOM here because this is the only safe context without
 * any locks held but let the oom killer triggered from the allocation context care
 * about the global OOM.
 */&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;


&lt;p&gt;The actual distinction is whether the &lt;code&gt;memcg&lt;/code&gt; field of &lt;code&gt;struct oom_control&lt;/code&gt; is NULL. If it holds a value, the victim is chosen from within that cgroup; if NULL, from the whole node (&lt;a href="https://github.com/torvalds/linux/blob/v6.12/include/linux/oom.h#L36-L37" rel="noopener noreferrer"&gt;oom.h&lt;/a&gt;). &lt;code&gt;is_memcg_oom()&lt;/code&gt; just checks this field.&lt;/p&gt;


&lt;pre class="highlight c"&gt;&lt;code&gt;  &lt;span class="cm"&gt;/* Memory cgroup in which oom is invoked, or NULL for global oom */&lt;/span&gt;
  &lt;span class="k"&gt;struct&lt;/span&gt; &lt;span class="n"&gt;mem_cgroup&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;memcg&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;


&lt;p&gt;&lt;strong&gt;&lt;code&gt;memcg&lt;/code&gt; stands for memory cgroup (memory control group)&lt;/strong&gt;, a short form of &lt;code&gt;struct mem_cgroup&lt;/code&gt; in the code above. Variable and function names in the source (&lt;code&gt;is_memcg_oom()&lt;/code&gt;, &lt;code&gt;try_charge_memcg()&lt;/code&gt;) use the same spelling, and the kernel documentation also writes "Memory cgroup (memcg)" (&lt;code&gt;Documentation/mm/hmm.rst&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;Note that the term &lt;code&gt;global OOM&lt;/code&gt; appears &lt;strong&gt;only in source comments and OOM logs&lt;/strong&gt;; it is not in the kernel documentation. &lt;code&gt;memcg OOM&lt;/code&gt; does appear as "memcg OOM killer" in &lt;code&gt;Documentation/admin-guide/cgroup-v1/memory.rst&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This is how they appear in the kernel log:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Name in this article&lt;/th&gt;
&lt;th&gt;How it appears in the kernel log&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;memcg OOM&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;constraint=CONSTRAINT_MEMCG&lt;/code&gt;, &lt;code&gt;oom_memcg=&amp;lt;cgroup path&amp;gt;&lt;/code&gt;, &lt;code&gt;Memory cgroup out of memory: Killed process ...&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;global OOM&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;constraint=CONSTRAINT_NONE&lt;/code&gt;, &lt;code&gt;,global_oom&lt;/code&gt;, &lt;code&gt;Out of memory: Killed process ...&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The prefix of the &lt;code&gt;Killed process&lt;/code&gt; line directly indicates the path (&lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/oom_kill.c#L1173-L1174" rel="noopener noreferrer"&gt;oom_kill.c&lt;/a&gt;):&lt;/p&gt;


&lt;pre class="highlight c"&gt;&lt;code&gt;      &lt;span class="n"&gt;oom_kill_process&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;oc&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="n"&gt;is_memcg_oom&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;oc&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;?&lt;/span&gt; &lt;span class="s"&gt;"Out of memory"&lt;/span&gt; &lt;span class="o"&gt;:&lt;/span&gt;
               &lt;span class="s"&gt;"Memory cgroup out of memory"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/blockquote&gt;

&lt;p&gt;Also note that &lt;strong&gt;eviction is a separate mechanism from an OOM kill&lt;/strong&gt;. It is the kubelet, not the kernel, that removes the Pod, and it is observed as &lt;code&gt;Evicted&lt;/code&gt;. It is covered below as well.&lt;/p&gt;
&lt;h2&gt;
  
  
  Quick reference: contradictions explained and a metric map
&lt;/h2&gt;

&lt;p&gt;This section expands the TL;DR: it explains the apparent contradictions from the introduction and maps each metric to the cgroup v2 value behind it.&lt;/p&gt;

&lt;p&gt;With that, the seemingly contradictory situations from the introduction can be explained as follows:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The working set has not reached the limit, yet &lt;code&gt;OOMKilled&lt;/code&gt;&lt;/strong&gt;: memcg OOM compares &lt;code&gt;memory.current&lt;/code&gt;, not the working set, against the limit (and merely reaching the limit does not trigger it; the kill happens when reclaim fails). Also, even if your container is under its limit, it can be killed in a node-wide memory exhaustion (global OOM). See My working set is below the limit, but the container was OOMKilled.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Node utilization computed from &lt;code&gt;MemAvailable&lt;/code&gt; is low, yet a Pod was evicted&lt;/strong&gt;: &lt;code&gt;MemAvailable&lt;/code&gt; counts part of the page cache as free, but the kubelet is working-set based and treats all active page cache as "in use". See Node usage computed from MemAvailable is low, but Pods are evicted.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The application does not leak memory, yet the working set keeps growing&lt;/strong&gt;: the working set includes active page cache. Page cache is not reclaimed until memory gets tight, so workloads with heavy file I/O or logging look like they keep growing. See Memory working set keeps growing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The RSS the application reports does not match &lt;code&gt;container_memory_rss&lt;/code&gt;&lt;/strong&gt;: the former is per-process RSS (including file mappings), and the latter is the anonymous pages charged to the container's cgroup. See The RSS my application reports differs from the RSS metric.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The mapping between these values and the metrics you see in &lt;code&gt;kubectl top&lt;/code&gt; and dashboards is as follows. &lt;strong&gt;Even among "memory usage" metrics, each one measures something different and corresponds to a different cgroup v2 field.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Emitted by&lt;/th&gt;
&lt;th&gt;Backing value in cgroup v2 / &lt;code&gt;/proc&lt;/code&gt;
&lt;/th&gt;
&lt;th&gt;Page cache handling&lt;/th&gt;
&lt;th&gt;Main use&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Container&lt;/td&gt;
&lt;td&gt;&lt;code&gt;container_memory_working_set_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;memory.current − inactive_file&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Active pages count as in use&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;kubectl top pod&lt;/code&gt;, usage visualization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Container&lt;/td&gt;
&lt;td&gt;&lt;code&gt;container_memory_rss&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;anon&lt;/code&gt; in &lt;code&gt;memory.stat&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Not included (anonymous pages only)&lt;/td&gt;
&lt;td&gt;Understanding what the app allocated&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Container&lt;/td&gt;
&lt;td&gt;&lt;code&gt;container_memory_usage_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;memory.current&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Everything included&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;What memcg OOM is evaluated against&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Node&lt;/td&gt;
&lt;td&gt;&lt;code&gt;container_memory_working_set_bytes{id="/"}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;anon + file − inactive_file&lt;/code&gt; in the root cgroup's &lt;code&gt;memory.stat&lt;/code&gt; (see below)&lt;/td&gt;
&lt;td&gt;Active pages count as in use&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Eviction decisions&lt;/strong&gt;, &lt;code&gt;kubectl top node&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Node&lt;/td&gt;
&lt;td&gt;&lt;code&gt;node_memory_MemAvailable_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;node-exporter&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;MemAvailable&lt;/code&gt; in &lt;code&gt;/proc/meminfo&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Reclaimable pages count as free&lt;/td&gt;
&lt;td&gt;Visualizing node memory utilization&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Node&lt;/td&gt;
&lt;td&gt;&lt;code&gt;kube_node_status_capacity{resource="memory"}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;kube-state-metrics&lt;/td&gt;
&lt;td&gt;(&lt;code&gt;status.capacity&lt;/code&gt; of the Node object)&lt;/td&gt;
&lt;td&gt;(capacity, not usage)&lt;/td&gt;
&lt;td&gt;Denominator when converting to utilization&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Looking at the same node, utilization computed from &lt;code&gt;node_memory_MemAvailable_bytes&lt;/code&gt; does not match the utilization the kubelet uses for eviction decisions. With workloads that do a lot of file I/O or logging, the difference can reach several GiB.&lt;/p&gt;

&lt;p&gt;Different emitters also mean &lt;strong&gt;different labels on the metrics&lt;/strong&gt;. cAdvisor metrics carry no label identifying the node, so when you compute per-node utilization you need to align labels with &lt;code&gt;kubelet_node_name&lt;/code&gt;, as shown later.&lt;/p&gt;

&lt;p&gt;Finally, the conclusions for diagnosis and remedy:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;To tell which OOM path occurred, check whether the Node has a &lt;code&gt;SystemOOM&lt;/code&gt; event.&lt;/strong&gt; If it does, it is global OOM, but &lt;strong&gt;its absence is not evidence of memcg OOM&lt;/strong&gt; (neither &lt;code&gt;OOMKilled&lt;/code&gt; nor metrics distinguish the paths). See Telling which path occurred.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The most effective remedy is to set the memory request to match actual usage.&lt;/strong&gt; It affects placement, eviction candidate selection and &lt;code&gt;oom_score_adj&lt;/code&gt; at the same time. PromQL for checking is collected in What to do about it.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  Kubernetes memory usage and cgroup v2
&lt;/h2&gt;

&lt;p&gt;This section sorts out which cgroup v2 value each metric corresponds to, and which value triggers OOM kills and evictions.&lt;/p&gt;
&lt;h3&gt;
  
  
  Memory breakdown in cgroup v2
&lt;/h3&gt;

&lt;p&gt;Container memory usage is managed in a per-container cgroup. In cgroup v2, &lt;code&gt;memory.current&lt;/code&gt; is the current usage and &lt;code&gt;memory.stat&lt;/code&gt; is its breakdown.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Example values in a container's cgroup (checked on the node)&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; /sys/fs/cgroup/&amp;lt;container cgroup&amp;gt;/memory.current
5368709120

&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; /sys/fs/cgroup/&amp;lt;container cgroup&amp;gt;/memory.stat
anon 2147483648           &lt;span class="c"&gt;# anonymous pages (heap, stack: memory not backed by a file)&lt;/span&gt;
file 3087007744           &lt;span class="c"&gt;# page cache (including tmpfs and shared memory)&lt;/span&gt;
shmem 268435456           &lt;span class="c"&gt;# the part of file that is tmpfs or shared memory&lt;/span&gt;
...
inactive_anon 134217728    &lt;span class="c"&gt;# anon LRU: not accessed for a while&lt;/span&gt;
active_anon 2013265920     &lt;span class="c"&gt;# anon LRU: recently accessed&lt;/span&gt;
inactive_file 1073741824   &lt;span class="c"&gt;# file LRU: not accessed for a while&lt;/span&gt;
active_file 1744830464     &lt;span class="c"&gt;# file LRU: recently accessed&lt;/span&gt;
unevictable 268435456      &lt;span class="c"&gt;# pages that cannot be reclaimed (here, the tmpfs part)&lt;/span&gt;
...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The following relationships hold in this example. &lt;strong&gt;Note that &lt;code&gt;inactive_file + active_file&lt;/code&gt; does not equal &lt;code&gt;file&lt;/code&gt;; it equals &lt;code&gt;file - shmem&lt;/code&gt;.&lt;/strong&gt; The next section explains why.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;anon + file + kernel memory &lt;span class="o"&gt;=&lt;/span&gt; 5368709120 &lt;span class="o"&gt;(&lt;/span&gt;memory.current&lt;span class="o"&gt;)&lt;/span&gt;
inactive_file + active_file &lt;span class="o"&gt;=&lt;/span&gt; 2818572288 &lt;span class="o"&gt;=&lt;/span&gt; file - shmem
inactive_anon + active_anon &lt;span class="o"&gt;=&lt;/span&gt; 2147483648 &lt;span class="o"&gt;=&lt;/span&gt; anon        &lt;span class="o"&gt;(&lt;/span&gt;shmem is not &lt;span class="k"&gt;in &lt;/span&gt;here&lt;span class="o"&gt;)&lt;/span&gt;
unevictable                 &lt;span class="o"&gt;=&lt;/span&gt;  268435456 &lt;span class="o"&gt;=&lt;/span&gt; shmem       &lt;span class="o"&gt;(&lt;/span&gt;tmpfs without swap&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The main fields of &lt;code&gt;memory.stat&lt;/code&gt; are:&lt;/p&gt;
&lt;h4&gt;
  
  
  &lt;code&gt;anon&lt;/code&gt; (anonymous pages)
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Memory not backed by a file. Regions the application allocated with &lt;code&gt;malloc&lt;/code&gt; and similar belong here.&lt;/li&gt;
&lt;li&gt;Unlike page cache, &lt;strong&gt;the kernel cannot reclaim it independently of the process&lt;/strong&gt;. With swap it can write pages out and free up space, but on nodes that do not use swap it is &lt;strong&gt;memory the kernel cannot reclaim&lt;/strong&gt; under pressure. When swap is unavailable, the kernel sets &lt;code&gt;scan_balance = SCAN_FILE&lt;/code&gt; and excludes the anon LRU from scanning (&lt;code&gt;get_scan_count()&lt;/code&gt; in &lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/vmscan.c#L2382-L2386" rel="noopener noreferrer"&gt;mm/vmscan.c&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;On the other hand, this value drops if the process itself returns memory to the OS with &lt;code&gt;munmap&lt;/code&gt; or &lt;code&gt;madvise(MADV_DONTNEED)&lt;/code&gt;. Note that calling &lt;code&gt;free&lt;/code&gt; does not reduce it unless the allocator or language runtime returns the memory to the OS.&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;
  
  
  &lt;code&gt;file&lt;/code&gt; (page cache)
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Contents of files held in memory by the kernel during reads and writes. The application did not allocate it explicitly.&lt;/li&gt;
&lt;li&gt;When memory runs short the kernel discards (reclaims) it automatically, so it is &lt;strong&gt;basically reclaimable memory&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Cache of files on disk is split into &lt;code&gt;active_file&lt;/code&gt; (recently accessed) and &lt;code&gt;inactive_file&lt;/code&gt; (not accessed for a while). Reclaim takes &lt;code&gt;inactive_file&lt;/code&gt; first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;active_file + inactive_file&lt;/code&gt; does not equal &lt;code&gt;file&lt;/code&gt;.&lt;/strong&gt; These are totals of the LRU lists, and as described below &lt;code&gt;shmem&lt;/code&gt; is on the anon LRU rather than the file LRU. Besides &lt;code&gt;shmem&lt;/code&gt;, pages that fall off the LRU, such as &lt;code&gt;mlock&lt;/code&gt;ed pages, also show up as a difference.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It includes not only the cache of files on disk but also &lt;code&gt;shmem&lt;/code&gt; (tmpfs and shared memory)&lt;/strong&gt;, because Linux treats these as page cache too. &lt;code&gt;shmem&lt;/code&gt; is written out to swap rather than to a file, so on nodes without swap it cannot be reclaimed. The breakdown is visible as &lt;code&gt;shmem&lt;/code&gt; in &lt;code&gt;memory.stat&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Warning: &lt;code&gt;shmem&lt;/code&gt; is counted in &lt;code&gt;file&lt;/code&gt;, but it is not on the file LRU.&lt;/strong&gt; The kernel classifies swap-backed pages as anon, and tmpfs pages are among them. &lt;code&gt;folio_is_file_lru()&lt;/code&gt; in &lt;a href="https://github.com/torvalds/linux/blob/v6.12/include/linux/mm_inline.h#L13-L31" rel="noopener noreferrer"&gt;mm_inline.h&lt;/a&gt; is defined as "0 if &lt;code&gt;@folio&lt;/code&gt; is a normal anonymous folio, &lt;strong&gt;a tmpfs folio&lt;/strong&gt; or otherwise ram or swap backed folio".&lt;/p&gt;

&lt;p&gt;The counters they are charged to are different, though. A tmpfs page is added to both &lt;code&gt;NR_FILE_PAGES&lt;/code&gt; (= &lt;code&gt;file&lt;/code&gt;) and &lt;code&gt;NR_SHMEM&lt;/code&gt; (= &lt;code&gt;shmem&lt;/code&gt;) (&lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/shmem.c#L818-L819" rel="noopener noreferrer"&gt;mm/shmem.c&lt;/a&gt;):&lt;/p&gt;


&lt;pre class="highlight c"&gt;&lt;code&gt;  &lt;span class="n"&gt;__lruvec_stat_mod_folio&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;folio&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;NR_FILE_PAGES&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;nr&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="n"&gt;__lruvec_stat_mod_folio&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;folio&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;NR_SHMEM&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;nr&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;


&lt;p&gt;Furthermore, for a Kubernetes &lt;code&gt;emptyDir&lt;/code&gt; with &lt;code&gt;medium: Memory&lt;/code&gt;, &lt;strong&gt;the kubelet mounts tmpfs with the &lt;code&gt;noswap&lt;/code&gt; option&lt;/strong&gt; (&lt;code&gt;generateTmpfsMountOptions()&lt;/code&gt; in &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/volume/emptydir/empty_dir.go#L601-L612" rel="noopener noreferrer"&gt;empty_dir.go&lt;/a&gt;). The kernel makes the inode of a &lt;code&gt;noswap&lt;/code&gt; tmpfs unevictable (&lt;code&gt;shmem_get_inode()&lt;/code&gt; in &lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/shmem.c#L2824-L2825" rel="noopener noreferrer"&gt;mm/shmem.c&lt;/a&gt;), so these pages sit on the &lt;strong&gt;unevictable LRU&lt;/strong&gt;, not the anon LRU.&lt;/p&gt;


&lt;pre class="highlight c"&gt;&lt;code&gt;  &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;sbinfo&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;noswap&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
      &lt;span class="n"&gt;mapping_set_unevictable&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;inode&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;i_mapping&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;


&lt;p&gt;So &lt;code&gt;shmem&lt;/code&gt; is accounted as follows:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;mount&lt;/th&gt;
&lt;th&gt;LRU&lt;/th&gt;
&lt;th&gt;Where it shows up in &lt;code&gt;memory.stat&lt;/code&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;emptyDir&lt;/code&gt; with &lt;code&gt;medium: Memory&lt;/code&gt; (&lt;code&gt;noswap&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;unevictable&lt;/td&gt;
&lt;td&gt;&lt;code&gt;unevictable&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Any other tmpfs / shared memory&lt;/td&gt;
&lt;td&gt;anon LRU&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;inactive_anon&lt;/code&gt; / &lt;code&gt;active_anon&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;In neither case does it enter &lt;code&gt;inactive_file&lt;/code&gt; / &lt;code&gt;active_file&lt;/code&gt;.&lt;/strong&gt; That matters in two ways later:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;inactive_file + active_file&lt;/code&gt; equals &lt;code&gt;file - shmem&lt;/code&gt;, not &lt;code&gt;file&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Only &lt;code&gt;inactive_file&lt;/code&gt; is subtracted from the working set, so &lt;strong&gt;&lt;code&gt;shmem&lt;/code&gt; stays in the working set in full&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;With workloads that write a lot of logs or read and write large files, the page cache can grow to several GiB even though the application did not intend that.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; &lt;code&gt;memory.current&lt;/code&gt; includes memory the kernel itself uses (slab, socket buffers and so on) in addition to &lt;code&gt;anon&lt;/code&gt; and &lt;code&gt;file&lt;/code&gt;. So adding &lt;code&gt;anon&lt;/code&gt; and &lt;code&gt;file&lt;/code&gt; alone does not give &lt;code&gt;memory.current&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;
  
  
  How cgroup v2 memory controls map to Pod settings
&lt;/h3&gt;
&lt;h4&gt;
  
  
  cgroup v2 memory control interface
&lt;/h4&gt;

&lt;p&gt;cgroup v2 offers four levels of control over memory usage. &lt;code&gt;memory.min&lt;/code&gt; and &lt;code&gt;memory.low&lt;/code&gt; are &lt;strong&gt;protections&lt;/strong&gt; ("do not reclaim up to here"), and &lt;code&gt;memory.high&lt;/code&gt; and &lt;code&gt;memory.max&lt;/code&gt; are &lt;strong&gt;limits&lt;/strong&gt; ("do not allow use beyond here").&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;Behavior&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;memory.min&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Protection (hard)&lt;/td&gt;
&lt;td&gt;Memory up to this value is not reclaimed. If nothing else can be reclaimed and usage cannot fit, an OOM kill occurs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;memory.low&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Protection (soft)&lt;/td&gt;
&lt;td&gt;Best-effort avoidance of reclaim up to this value. It is reclaimed if nothing else can be&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;memory.high&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Limit (soft)&lt;/td&gt;
&lt;td&gt;Exceeding it applies heavy reclaim pressure and throttles the cgroup's processes. No OOM kill&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;memory.max&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Limit (hard)&lt;/td&gt;
&lt;td&gt;Exceeding it triggers reclaim, and if usage still does not fit, the cgroup's OOM killer runs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Other files that report state:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;memory.current&lt;/code&gt;: current usage (what &lt;code&gt;container_memory_usage_bytes&lt;/code&gt; is backed by)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;memory.stat&lt;/code&gt;: usage breakdown&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;memory.events&lt;/code&gt;: counts of &lt;code&gt;low&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt;, &lt;code&gt;max&lt;/code&gt;, &lt;code&gt;oom&lt;/code&gt;, &lt;code&gt;oom_kill&lt;/code&gt; and &lt;code&gt;oom_group_kill&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;memory.swap.max&lt;/code&gt;: swap limit&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;memory.pressure&lt;/code&gt;: PSI for waiting on memory allocation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For exact definitions of each file see the kernel's &lt;a href="https://docs.kernel.org/admin-guide/cgroup-v2.html#memory-interface-files" rel="noopener noreferrer"&gt;Memory Interface Files&lt;/a&gt;, and for how Kubernetes treats cgroup v2 see &lt;a href="https://kubernetes.io/docs/concepts/architecture/cgroups/" rel="noopener noreferrer"&gt;About cgroup v2&lt;/a&gt;.&lt;/p&gt;
&lt;h4&gt;
  
  
  How Pod settings are reflected
&lt;/h4&gt;

&lt;p&gt;Pod settings are reflected into cgroup v2 as follows.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pod setting&lt;/th&gt;
&lt;th&gt;Target in cgroup v2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;resources.limits.memory&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;memory.max&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;resources.requests.memory&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Not reflected&lt;/strong&gt; (with the default kubelet configuration in v1.37; you can reflect it by changing settings → KEP-2570)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;resources.limits.cpu&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;cpu.max&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;resources.requests.cpu&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;cpu.weight&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;cgroup v2 is a directory tree. Its top is the &lt;strong&gt;root cgroup&lt;/strong&gt; (&lt;code&gt;/sys/fs/cgroup&lt;/code&gt;), where any process that is in no child cgroup belongs, so &lt;strong&gt;it contains not only Pods but every process on the node, such as the kubelet and containerd&lt;/strong&gt;. Kubernetes creates the hierarchy below it: "all Pods on the node → Pod → container".&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# With the systemd cgroup driver (the kubeadm default; names differ with the cgroupfs driver, see below)&lt;/span&gt;
/sys/fs/cgroup                                  ← root cgroup &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
│                                                 contains every process on the node
│                                                 the value used &lt;span class="k"&gt;for &lt;/span&gt;eviction decisions
│
├── system.slice/                               kubelet, containerd, sshd, etc.
│   ├── kubelet.service                           &lt;span class="o"&gt;(&lt;/span&gt;processes that are not Pods&lt;span class="o"&gt;)&lt;/span&gt;
│   └── containerd.service
│
└── kubepods.slice/                             ← all Pods on the node
    │                                             &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/kubepods.slice"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
    ├── kubepods-pod&amp;lt;UID&amp;gt;.slice/                ← Pod &lt;span class="o"&gt;(&lt;/span&gt;Guaranteed&lt;span class="o"&gt;)&lt;/span&gt;
    │   ├── cri-containerd-&amp;lt;ID&amp;gt;.scope           ← container
    │   └── cri-containerd-&amp;lt;pause ID&amp;gt;.scope     ← pause container
    │
    ├── kubepods-burstable.slice/                 &lt;span class="o"&gt;(&lt;/span&gt;intermediate slice per QoS class&lt;span class="o"&gt;)&lt;/span&gt;
    │   └── kubepods-burstable-pod&amp;lt;UID&amp;gt;.slice/  ← Pod &lt;span class="o"&gt;(&lt;/span&gt;Burstable&lt;span class="o"&gt;)&lt;/span&gt;
    │       └── cri-containerd-&amp;lt;ID&amp;gt;.scope       ← container
    │
    └── kubepods-besteffort.slice/
        └── kubepods-besteffort-pod&amp;lt;UID&amp;gt;.slice/ ← Pod &lt;span class="o"&gt;(&lt;/span&gt;BestEffort&lt;span class="o"&gt;)&lt;/span&gt;
            └── cri-containerd-&amp;lt;ID&amp;gt;.scope       ← container
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;When filtering metrics, the root cgroup is &lt;code&gt;id="/"&lt;/code&gt;, all Pods on the node are &lt;code&gt;id="/kubepods.slice"&lt;/code&gt;, a Pod is &lt;code&gt;container="", image="", pod!=""&lt;/code&gt;, and a container is &lt;code&gt;container!="", image!=""&lt;/code&gt;. &lt;strong&gt;What the kubelet uses for eviction decisions is the working set of the topmost root cgroup&lt;/strong&gt; (see below). That is a different level from memcg OOM caused by a container or Pod exceeding its limit.&lt;/p&gt;

&lt;p&gt;Memory limits are set not only on containers but also on higher levels.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;cgroup&lt;/th&gt;
&lt;th&gt;Memory limit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Whole node (root cgroup)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Not set&lt;/strong&gt; (the root has no &lt;code&gt;memory.max&lt;/code&gt; → see below)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;All Pods on the node&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/kubepods.slice&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;Capacity - kube-reserved - system-reserved&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pod&lt;/td&gt;
&lt;td&gt;Differs by QoS class (see below)&lt;/td&gt;
&lt;td&gt;The Pod's effective limit (see below). Set only when every container has a limit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Container&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;cri-containerd-&amp;lt;ID&amp;gt;.scope&lt;/code&gt; under the Pod's cgroup&lt;/td&gt;
&lt;td&gt;That container's &lt;code&gt;limits.memory&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A Pod's effective limit is not simply the sum of its regular containers' &lt;code&gt;limits.memory&lt;/code&gt;. If there are init containers, their maximum is also considered, and &lt;code&gt;spec.overhead&lt;/code&gt; is added if set. For the calculation see &lt;a href="https://kubernetes.io/docs/concepts/workloads/pods/init-containers/#resource-sharing-within-containers" rel="noopener noreferrer"&gt;Resource sharing within containers&lt;/a&gt; and &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/pod-overhead/" rel="noopener noreferrer"&gt;Pod Overhead&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;So &lt;strong&gt;even if a container has not reached its own limit, an OOM kill can occur when the Pod as a whole or &lt;code&gt;/kubepods.slice&lt;/code&gt; hits its limit&lt;/strong&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; The paths above assume the kubelet uses the systemd cgroup driver. With the cgroupfs driver, paths look like &lt;code&gt;/kubepods/burstable/pod&amp;lt;UID&amp;gt;/&amp;lt;container ID&amp;gt;&lt;/code&gt;, with no &lt;code&gt;.slice&lt;/code&gt; or &lt;code&gt;.scope&lt;/code&gt; and no &lt;code&gt;cri-containerd-&lt;/code&gt; prefix on the container.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The kubelet's own default is &lt;code&gt;cgroupfs&lt;/code&gt;, but kubeadm sets &lt;code&gt;systemd&lt;/code&gt; when the value is empty&lt;/strong&gt; (&lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/apis/config/v1beta1/defaults.go#L168-L170" rel="noopener noreferrer"&gt;defaults.go&lt;/a&gt;, &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/cmd/kubeadm/app/componentconfigs/kubelet.go#L194-L196" rel="noopener noreferrer"&gt;kubelet.go&lt;/a&gt;). You can check which one is in use with:&lt;/p&gt;


&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl get &lt;span class="nt"&gt;--raw&lt;/span&gt; &lt;span class="s2"&gt;"/api/v1/nodes/&amp;lt;node name&amp;gt;/proxy/configz"&lt;/span&gt; | jq &lt;span class="s1"&gt;'.kubeletconfig.cgroupDriver'&lt;/span&gt;
&lt;span class="s2"&gt;"systemd"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;


&lt;p&gt;The &lt;code&gt;id&lt;/code&gt; label of cAdvisor metrics also changes with the cgroup driver, so when filtering in PromQL, check the value in your own cluster with this query:&lt;/p&gt;


&lt;pre class="highlight shell"&gt;&lt;code&gt;group by &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;container_memory_working_set_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;~&lt;span class="s2"&gt;"/kubepods|/kubepods.slice"&lt;/span&gt;&lt;span class="o"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;


&lt;p&gt;A Pod's cgroup path depends on its QoS class. As the tree above shows, Burstable and BestEffort go through an intermediate slice per QoS class.&lt;/p&gt;

&lt;p&gt;In the &lt;code&gt;&amp;lt;UID&amp;gt;&lt;/code&gt; part, systemd escaping replaces the hyphens of the Pod UID with underscores (example: &lt;code&gt;6983351e-4ede-...&lt;/code&gt; → &lt;code&gt;pod6983351e_4ede_...&lt;/code&gt;). The UID of a static Pod (mirror Pod) is a hash with no hyphens, so it looks like &lt;code&gt;pod1ff70b81885169da4e9e32c9530b77c9.slice&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The QoS class and UID of a Pod can be read from the Pod resource, and once you have those two you can assemble the cgroup path.&lt;/p&gt;


&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl get pod &lt;span class="nt"&gt;-n&lt;/span&gt; &amp;lt;namespace&amp;gt; &amp;lt;Pod name&amp;gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;jsonpath&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'{.status.qosClass}{"\n"}{.metadata.uid}{"\n"}'&lt;/span&gt;
Burstable
6983351e-4ede-4611-8bab-b986614af18d

&lt;span class="c"&gt;# To list them&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl get pods &lt;span class="nt"&gt;-n&lt;/span&gt; &amp;lt;namespace&amp;gt; &lt;span class="nt"&gt;-o&lt;/span&gt; custom-columns&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'NAME:.metadata.name,QOS:.status.qosClass'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/blockquote&gt;
&lt;h3&gt;
  
  
  Container metrics
&lt;/h3&gt;

&lt;p&gt;The &lt;a href="https://github.com/google/cadvisor" rel="noopener noreferrer"&gt;cAdvisor&lt;/a&gt; built into the kubelet exposes the cgroup values above as Prometheus metrics. They are served from the kubelet's &lt;code&gt;/metrics/cadvisor&lt;/code&gt; endpoint, and you do not need to run a separate cAdvisor Pod or DaemonSet.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Emitted by&lt;/th&gt;
&lt;th&gt;Backing value in cgroup v2&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_usage_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;memory.current&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Usage including anonymous pages, page cache and kernel memory. The value memcg OOM is evaluated against&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_working_set_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;memory.current - inactive_file&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Memory considered "not immediately reclaimable". The value Kubernetes treats as usage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_rss&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;anon&lt;/code&gt; in &lt;code&gt;memory.stat&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Amount of anonymous pages. Unlike true RSS, it does not include &lt;code&gt;mmap&lt;/code&gt;ed file pages (see below)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_cache&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;file&lt;/code&gt; in &lt;code&gt;memory.stat&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Page cache size. Includes &lt;code&gt;shmem&lt;/code&gt; (tmpfs and shared memory)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_total_active_file_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;active_file&lt;/code&gt; in &lt;code&gt;memory.stat&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Active page cache&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_total_inactive_file_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;inactive_file&lt;/code&gt; in &lt;code&gt;memory.stat&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Inactive page cache&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_mapped_file&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;file_mapped&lt;/code&gt; in &lt;code&gt;memory.stat&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Amount of &lt;code&gt;mmap&lt;/code&gt;ed files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_pressure_memory_waiting_seconds_total&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;total&lt;/code&gt; of &lt;code&gt;some&lt;/code&gt; in &lt;code&gt;memory.pressure&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Time some processes were waiting for memory (see below)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_pressure_memory_stalled_seconds_total&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;total&lt;/code&gt; of &lt;code&gt;full&lt;/code&gt; in &lt;code&gt;memory.pressure&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Time all processes were stalled waiting for memory (see below)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_oom_events_total&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;kubelet&lt;/strong&gt; (v1.37; the name starts with &lt;code&gt;container_&lt;/code&gt; but it is not from cAdvisor)&lt;/td&gt;
&lt;td&gt;(not a cgroup value)&lt;/td&gt;
&lt;td&gt;Counted by parsing &lt;code&gt;oom-kill:&lt;/code&gt; lines in the kernel log. &lt;strong&gt;When the container terminates its cgroup disappears and the series is lost&lt;/strong&gt; (whether or not it is recreated)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; The kubelet has endpoints such as &lt;code&gt;/metrics&lt;/code&gt;, &lt;code&gt;/metrics/cadvisor&lt;/code&gt; and &lt;code&gt;/metrics/resource&lt;/code&gt; (&lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/server/server.go#L110-L113" rel="noopener noreferrer"&gt;server.go&lt;/a&gt;). The last two relate to container and node usage. &lt;code&gt;/metrics/cadvisor&lt;/code&gt; exposes the &lt;code&gt;container_*&lt;/code&gt; metrics above, while &lt;code&gt;/metrics/resource&lt;/code&gt; exposes &lt;strong&gt;only CPU, memory and swap usage&lt;/strong&gt; (&lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/metrics/collectors/resource_metrics.go#L29-L120" rel="noopener noreferrer"&gt;resource_metrics.go&lt;/a&gt;). The latter has &lt;code&gt;container_memory_working_set_bytes&lt;/code&gt; and also &lt;strong&gt;per-Pod and per-node series (&lt;code&gt;pod_memory_working_set_bytes&lt;/code&gt;, &lt;code&gt;node_memory_working_set_bytes&lt;/code&gt;)&lt;/strong&gt; (see &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/metrics/collectors/resource_metrics.go#L86-L91" rel="noopener noreferrer"&gt;&lt;code&gt;podMemoryUsageDesc&lt;/code&gt;&lt;/a&gt; and &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/metrics/collectors/resource_metrics.go#L37-L42" rel="noopener noreferrer"&gt;&lt;code&gt;nodeMemoryUsageDesc&lt;/code&gt;&lt;/a&gt; in resource_metrics.go). When to use which for per-Pod usage is covered in Align memory request with actual usage.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The important thing here is the definition of the working set.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;container_memory_working_set_bytes
  &lt;span class="o"&gt;=&lt;/span&gt; container_memory_usage_bytes - container_memory_total_inactive_file_bytes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;In other words, &lt;strong&gt;only inactive page cache is subtracted from the working set, and active page cache stays in as "in use"&lt;/strong&gt;. For the implementation see cAdvisor's &lt;a href="https://github.com/google/cadvisor/blob/v0.57.0/container/libcontainer/handler.go#L858-L867" rel="noopener noreferrer"&gt;handler.go&lt;/a&gt;.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;    &lt;span class="n"&gt;workingSet&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;ret&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Memory&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Usage&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ok&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MemoryStats&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Stats&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;inactiveFileKeyName&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt; &lt;span class="n"&gt;ok&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;ret&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Memory&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;TotalInactiveFile&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;workingSet&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;workingSet&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="m"&gt;0&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;workingSet&lt;/span&gt; &lt;span class="o"&gt;-=&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="n"&gt;ret&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Memory&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;WorkingSet&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;workingSet&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;As a breakdown, it is roughly the sum of the following. RSS (&lt;code&gt;container_memory_rss&lt;/code&gt;) is exactly &lt;code&gt;anon&lt;/code&gt; in &lt;code&gt;memory.stat&lt;/code&gt; under cgroup v2, that is, the amount of anonymous pages, so it is included in the working set.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;container_memory_working_set_bytes
  ≒ container_memory_rss                        &lt;span class="o"&gt;(&lt;/span&gt;anonymous pages&lt;span class="p"&gt;;&lt;/span&gt; shmem not included&lt;span class="o"&gt;)&lt;/span&gt;
  + shmem                                       &lt;span class="o"&gt;(&lt;/span&gt;tmpfs and shared memory&lt;span class="p"&gt;;&lt;/span&gt; stays &lt;span class="k"&gt;in &lt;/span&gt;full&lt;span class="o"&gt;)&lt;/span&gt;
  + container_memory_total_active_file_bytes    &lt;span class="o"&gt;(&lt;/span&gt;active page cache&lt;span class="o"&gt;)&lt;/span&gt;
  + unreclaimable pages, kernel memory, etc.    &lt;span class="o"&gt;(&lt;/span&gt;unevictable, slab, socket buffers, ...&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;code&gt;shmem&lt;/code&gt; is listed as a separate term because &lt;strong&gt;neither &lt;code&gt;container_memory_rss&lt;/code&gt; (= &lt;code&gt;anon&lt;/code&gt; = &lt;code&gt;NR_ANON_MAPPED&lt;/code&gt;) nor &lt;code&gt;active_file&lt;/code&gt; includes &lt;code&gt;shmem&lt;/code&gt;&lt;/strong&gt;. As described above, &lt;code&gt;shmem&lt;/code&gt; is not on the file LRU, so it never appears in &lt;code&gt;active_file&lt;/code&gt; / &lt;code&gt;inactive_file&lt;/code&gt;, and since only &lt;code&gt;inactive_file&lt;/code&gt; is subtracted from the working set, it remains in full. No &lt;code&gt;container_*&lt;/code&gt; metric is exposed for &lt;code&gt;shmem&lt;/code&gt;, so this term can only be checked from &lt;code&gt;memory.stat&lt;/code&gt; on the node.&lt;/p&gt;

&lt;p&gt;It is &lt;code&gt;≒&lt;/code&gt; because, besides &lt;code&gt;shmem&lt;/code&gt;, there are also no exposed metrics for the kernel memory and locked pages contained in &lt;code&gt;memory.current&lt;/code&gt;. For an exact breakdown, read &lt;code&gt;memory.stat&lt;/code&gt; on the node directly.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Caution: RSS in &lt;code&gt;container_memory_rss&lt;/code&gt; stands for Resident Set Size, but its scope differs from the real meaning.&lt;/strong&gt; True RSS means "the pages a process holds in physical memory" and includes &lt;code&gt;mmap&lt;/code&gt;ed file pages.&lt;/p&gt;

&lt;p&gt;But &lt;code&gt;memory.stat&lt;/code&gt; in cgroup v2 has no &lt;code&gt;rss&lt;/code&gt; field, and cAdvisor &lt;a href="https://github.com/google/cadvisor/blob/v0.57.0/container/libcontainer/handler.go#L819-L834" rel="noopener noreferrer"&gt;exposes &lt;code&gt;anon&lt;/code&gt; as &lt;code&gt;container_memory_rss&lt;/code&gt;&lt;/a&gt;. The kernel documentation defines &lt;code&gt;anon&lt;/code&gt; as "Amount of memory used in anonymous mappings such as brk(), sbrk(), and mmap(MAP_ANONYMOUS)", which does not include file mappings. So it does not match the &lt;code&gt;RSS&lt;/code&gt; column of &lt;code&gt;ps&lt;/code&gt; or &lt;code&gt;VmRSS&lt;/code&gt; in &lt;code&gt;/proc/[pid]/status&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Warning:&lt;/strong&gt; A large &lt;code&gt;container_memory_working_set_bytes&lt;/code&gt; does not necessarily mean the application needs that much memory. It may just be accumulated page cache. To see how much memory the application has actually allocated, compare against &lt;code&gt;container_memory_rss&lt;/code&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;
  
  
  Node metrics
&lt;/h3&gt;

&lt;p&gt;There are two families of metrics representing whole-node memory usage, with different emitters.&lt;/p&gt;
&lt;h4&gt;
  
  
  From cAdvisor (the kubelet's view)
&lt;/h4&gt;

&lt;p&gt;cAdvisor exposes memory usage not only for containers but also for the root cgroup (&lt;code&gt;id="/"&lt;/code&gt;, the topmost cgroup that contains every process on the node, as described above) and for the cgroup of all Pods. The &lt;code&gt;id&lt;/code&gt; of the all-Pods cgroup depends on the cgroup driver, so the queries below match both values with &lt;code&gt;=~&lt;/code&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;container_memory_working_set_bytes{id="/"}&lt;/code&gt;: working set of the whole node (Pods + other processes). &lt;strong&gt;But the root cgroup has no &lt;code&gt;memory.current&lt;/code&gt;, so the backing value differs from containers&lt;/strong&gt; (see below)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;container_memory_working_set_bytes{id=~"/kubepods|/kubepods.slice"}&lt;/code&gt;: working set of all Pods on the node&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;container_memory_total_active_file_bytes{id="/"}&lt;/code&gt;: active page cache of the whole node&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both the &lt;code&gt;MEMORY(bytes)&lt;/code&gt; column of &lt;code&gt;kubectl top node&lt;/code&gt; and the kubelet's eviction decision described below use this working set.&lt;/p&gt;
&lt;h4&gt;
  
  
  From node-exporter (&lt;code&gt;/proc/meminfo&lt;/code&gt; as is)
&lt;/h4&gt;

&lt;p&gt;&lt;a href="https://github.com/prometheus/node_exporter" rel="noopener noreferrer"&gt;node-exporter&lt;/a&gt; turns the contents of &lt;code&gt;/proc/meminfo&lt;/code&gt; directly into metrics.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;node_memory_MemTotal_bytes&lt;/code&gt;: total physical memory on the node&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;node_memory_MemAvailable_bytes&lt;/code&gt;: the kernel's estimate of "how much can be allocated when new memory is requested"&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;node_memory_Cached_bytes&lt;/code&gt;: page cache size. As the kernel documentation says, "In-memory cache for files read from the disk (the pagecache) as well as tmpfs &amp;amp; shmem", it includes tmpfs and shared memory. The breakdown is in &lt;code&gt;node_memory_Shmem_bytes&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;code&gt;node_memory_MemAvailable_bytes&lt;/code&gt; &lt;strong&gt;counts reclaimable page cache as free&lt;/strong&gt;. Computing node memory utilization as &lt;code&gt;1 - MemAvailable / MemTotal&lt;/code&gt; is also in node-exporter's official mixin as the &lt;code&gt;NodeMemoryHighUtilization&lt;/code&gt; alert (&lt;a href="https://github.com/prometheus/node_exporter/blob/v1.12.1/docs/node-mixin/alerts/alerts.libsonnet#L364-L367" rel="noopener noreferrer"&gt;alerts.libsonnet&lt;/a&gt;). Its &lt;code&gt;expr&lt;/code&gt; uses variables for the selector and threshold; stripped of those it is &lt;code&gt;100 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes * 100) &amp;gt; &amp;lt;threshold&amp;gt;&lt;/code&gt;. However, as shown below, this value does not match the utilization the kubelet uses for eviction decisions.&lt;/p&gt;
&lt;h3&gt;
  
  
  Values the kubelet uses for eviction decisions
&lt;/h3&gt;

&lt;p&gt;The kubelet checks the node's free memory (&lt;code&gt;memory.available&lt;/code&gt;) every 10 seconds and starts evicting Pods when it falls below the threshold. This &lt;code&gt;memory.available&lt;/code&gt; is not computed from &lt;code&gt;MemAvailable&lt;/code&gt; in &lt;code&gt;/proc/meminfo&lt;/code&gt;; &lt;strong&gt;it is computed from the root cgroup's working set&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;memory.available &lt;span class="o"&gt;=&lt;/span&gt; node_memory_MemTotal_bytes - container_memory_working_set_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;For the implementation see Kubernetes' &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/stats/helper.go#L63-L87" rel="noopener noreferrer"&gt;helper.go&lt;/a&gt; and &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/eviction/helpers_others.go#L27-L37" rel="noopener noreferrer"&gt;helpers_others.go&lt;/a&gt;. The 10-second interval is defined as &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/kubelet.go#L197" rel="noopener noreferrer"&gt;&lt;code&gt;evictionMonitoringPeriod&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Warning: The root cgroup working set used for eviction decisions (and by &lt;code&gt;kubectl top node&lt;/code&gt;) is not the same thing as a container's working set.&lt;/strong&gt; In cgroup v2, &lt;code&gt;memory.current&lt;/code&gt; has the &lt;code&gt;CFTYPE_NOT_ON_ROOT&lt;/code&gt; flag, so &lt;strong&gt;the root cgroup has no &lt;code&gt;memory.current&lt;/code&gt;&lt;/strong&gt; (&lt;code&gt;memory_files[]&lt;/code&gt; in &lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/memcontrol.c#L4350-L4354" rel="noopener noreferrer"&gt;mm/memcontrol.c&lt;/a&gt;).&lt;/p&gt;


&lt;pre class="highlight c"&gt;&lt;code&gt;  &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"current"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;flags&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;CFTYPE_NOT_ON_ROOT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;read_u64&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;memory_current_read&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;


&lt;p&gt;&lt;code&gt;memory.stat&lt;/code&gt;, on the other hand, has no such flag, so it can be read even on the root. The cgroup library cAdvisor uses therefore substitutes &lt;code&gt;anon + file&lt;/code&gt; from &lt;code&gt;memory.stat&lt;/code&gt; as the usage when reading &lt;code&gt;memory.current&lt;/code&gt; fails with &lt;code&gt;ENOENT&lt;/code&gt;, for the root only (&lt;code&gt;rootStatsFromMeminfo()&lt;/code&gt; in &lt;a href="https://github.com/opencontainers/cgroups/blob/v0.0.6/fs2/memory.go#L229-L231" rel="noopener noreferrer"&gt;opencontainers/cgroups&lt;/a&gt;).&lt;/p&gt;


&lt;pre class="highlight go"&gt;&lt;code&gt;  &lt;span class="c"&gt;// sum `anon` + `file` to report the same value as `usage_in_bytes` in v1.&lt;/span&gt;
  &lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MemoryStats&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Usage&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Usage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MemoryStats&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Stats&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"anon"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MemoryStats&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Stats&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;"file"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;


&lt;p&gt;So for the &lt;code&gt;id="/"&lt;/code&gt; series:&lt;/p&gt;


&lt;pre class="highlight shell"&gt;&lt;code&gt;container_memory_usage_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;       &lt;span class="o"&gt;=&lt;/span&gt; anon + file
container_memory_working_set_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; anon + file - inactive_file
&lt;/code&gt;&lt;/pre&gt;


&lt;p&gt;&lt;strong&gt;Unlike &lt;code&gt;memory.current&lt;/code&gt;, this includes no kernel memory at all, such as slab, &lt;code&gt;pagetables&lt;/code&gt;, &lt;code&gt;kernel_stack&lt;/code&gt;, &lt;code&gt;percpu&lt;/code&gt; and &lt;code&gt;sock&lt;/code&gt;.&lt;/strong&gt; This is mainly why the node working set tends to come out smaller than usage derived from &lt;code&gt;/proc/meminfo&lt;/code&gt;. For container and Pod cgroups &lt;code&gt;memory.current&lt;/code&gt; can be read directly, so this substitution does not occur.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Because the working set includes active page cache, &lt;strong&gt;the kubelet treats active page cache as "memory in use"&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The default threshold is only the hard eviction &lt;code&gt;memory.available&amp;lt;100Mi&lt;/code&gt;, as in &lt;code&gt;DefaultEvictionHard&lt;/code&gt; in &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/eviction/defaults_linux.go#L21-L28" rel="noopener noreferrer"&gt;&lt;code&gt;defaults_linux.go&lt;/code&gt;&lt;/a&gt;; no soft eviction is configured by default.&lt;/p&gt;

&lt;p&gt;In practice, though, I think it is better to override this value explicitly. In the v1.37 cluster I used for verification I also pass the following flags to the kubelet to set both hard and soft thresholds.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nt"&gt;--eviction-hard&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;memory.available&amp;lt;5%,nodefs.available&amp;lt;5%,pid.available&amp;lt;5%
&lt;span class="nt"&gt;--eviction-soft&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;memory.available&amp;lt;10%,nodefs.available&amp;lt;10%,pid.available&amp;lt;10%
&lt;span class="nt"&gt;--eviction-soft-grace-period&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;memory.available&lt;span class="o"&gt;=&lt;/span&gt;2m,nodefs.available&lt;span class="o"&gt;=&lt;/span&gt;5s,pid.available&lt;span class="o"&gt;=&lt;/span&gt;1m
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The default &lt;code&gt;100Mi&lt;/code&gt; is a fixed value, so &lt;strong&gt;it becomes relatively thinner on nodes with more memory&lt;/strong&gt;. If you specify percentages, the threshold follows the node size. Placing soft eviction ahead of hard (10 %) and giving it a grace period with &lt;code&gt;--eviction-soft-grace-period&lt;/code&gt; avoids evicting Pods on momentary spikes while still evicting them when the pressure persists.&lt;/p&gt;

&lt;p&gt;You can check your cluster's settings with the following query.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl get &lt;span class="nt"&gt;--raw&lt;/span&gt; &lt;span class="s2"&gt;"/api/v1/nodes/&amp;lt;node name&amp;gt;/proxy/configz"&lt;/span&gt; | jq &lt;span class="s1"&gt;'.kubeletconfig | {evictionHard, evictionSoft, evictionSoftGracePeriod}'&lt;/span&gt;
&lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="s2"&gt;"evictionHard"&lt;/span&gt;: &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="s2"&gt;"memory.available"&lt;/span&gt;: &lt;span class="s2"&gt;"5%"&lt;/span&gt;,
    &lt;span class="s2"&gt;"nodefs.available"&lt;/span&gt;: &lt;span class="s2"&gt;"5%"&lt;/span&gt;,
    &lt;span class="s2"&gt;"pid.available"&lt;/span&gt;: &lt;span class="s2"&gt;"5%"&lt;/span&gt;
  &lt;span class="o"&gt;}&lt;/span&gt;,
  &lt;span class="s2"&gt;"evictionSoft"&lt;/span&gt;: &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="s2"&gt;"memory.available"&lt;/span&gt;: &lt;span class="s2"&gt;"10%"&lt;/span&gt;,
    &lt;span class="s2"&gt;"nodefs.available"&lt;/span&gt;: &lt;span class="s2"&gt;"10%"&lt;/span&gt;,
    &lt;span class="s2"&gt;"pid.available"&lt;/span&gt;: &lt;span class="s2"&gt;"10%"&lt;/span&gt;
  &lt;span class="o"&gt;}&lt;/span&gt;,
  &lt;span class="s2"&gt;"evictionSoftGracePeriod"&lt;/span&gt;: &lt;span class="o"&gt;{&lt;/span&gt;
    &lt;span class="s2"&gt;"memory.available"&lt;/span&gt;: &lt;span class="s2"&gt;"2m"&lt;/span&gt;,
    &lt;span class="s2"&gt;"nodefs.available"&lt;/span&gt;: &lt;span class="s2"&gt;"5s"&lt;/span&gt;,
    &lt;span class="s2"&gt;"pid.available"&lt;/span&gt;: &lt;span class="s2"&gt;"1m"&lt;/span&gt;
  &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; The denominator is Capacity (&lt;code&gt;node.status.capacity.memory&lt;/code&gt;, &lt;code&gt;kube_node_status_capacity{resource="memory"}&lt;/code&gt;), not Allocatable. These are the same value as &lt;code&gt;node_memory_MemTotal_bytes&lt;/code&gt;. The kubelet uses cAdvisor's value as the machine's memory capacity, and cAdvisor reads it from &lt;code&gt;MemTotal&lt;/code&gt; in &lt;code&gt;/proc/meminfo&lt;/code&gt; (&lt;a href="https://github.com/google/cadvisor/blob/v0.57.0/container/raw/handler.go#L137-L146" rel="noopener noreferrer"&gt;raw/handler.go&lt;/a&gt;). A threshold like &lt;code&gt;memory.available&amp;lt;5%&lt;/code&gt; is also evaluated as a percentage of this Capacity.&lt;/p&gt;

&lt;p&gt;Allocatable is Capacity minus the system reservation and the hard eviction threshold, and is the upper bound the scheduler uses when assigning Pod requests. It is not used to decide whether eviction happens. See &lt;a href="https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/" rel="noopener noreferrer"&gt;Reserve Compute Resources for System Daemons&lt;/a&gt; for details.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;
  
  
  Why node-exporter and kubelet usage disagree
&lt;/h3&gt;

&lt;p&gt;Putting the above together, there are two views of the same node.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;View&lt;/th&gt;
&lt;th&gt;Calculation&lt;/th&gt;
&lt;th&gt;Emitted by&lt;/th&gt;
&lt;th&gt;Active page cache&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Node memory utilization from node-exporter&lt;/td&gt;
&lt;td&gt;&lt;code&gt;1 - MemAvailable / MemTotal&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;node-exporter (&lt;code&gt;/proc/meminfo&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;The part estimated reclaimable counts as free&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kubelet's eviction decision&lt;/td&gt;
&lt;td&gt;&lt;code&gt;node working set / MemTotal&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet; root cgroup &lt;code&gt;anon + file - inactive_file&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Counted entirely as in use&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;MemAvailable&lt;/code&gt; is the kernel's estimate of "how much can be newly allocated", and it adds part of the page cache and the reclaimable slab to free memory. It does not count the whole page cache as free. The implementation is &lt;code&gt;si_mem_available()&lt;/code&gt; in &lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/show_mem.c#L43-L67" rel="noopener noreferrer"&gt;mm/show_mem.c&lt;/a&gt;.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;    &lt;span class="cm"&gt;/*
     * Estimate the amount of memory available for userspace allocations,
     * without causing swapping or OOM.
     */&lt;/span&gt;
    &lt;span class="n"&gt;available&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;global_zone_page_state&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;NR_FREE_PAGES&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;totalreserve_pages&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="cm"&gt;/*
     * Not all the page cache can be freed, otherwise the system will
     * start swapping or thrashing. Assume at least half of the page
     * cache, or the low watermark worth of cache, needs to stay.
     */&lt;/span&gt;
    &lt;span class="n"&gt;pagecache&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;global_node_page_state&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;NR_ACTIVE_FILE&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
        &lt;span class="n"&gt;global_node_page_state&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;NR_INACTIVE_FILE&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="n"&gt;pagecache&lt;/span&gt; &lt;span class="o"&gt;-=&lt;/span&gt; &lt;span class="n"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pagecache&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;wmark_low&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="n"&gt;available&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;pagecache&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="cm"&gt;/*
     * Part of the reclaimable slab and other kernel memory consists of
     * items that are in use, and cannot be freed. Cap this estimate at the
     * low watermark.
     */&lt;/span&gt;
    &lt;span class="n"&gt;reclaimable&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;global_node_page_state_pages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;NR_SLAB_RECLAIMABLE_B&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
        &lt;span class="n"&gt;global_node_page_state&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;NR_KERNEL_MISC_RECLAIMABLE&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="n"&gt;reclaimable&lt;/span&gt; &lt;span class="o"&gt;-=&lt;/span&gt; &lt;span class="n"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;reclaimable&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;wmark_low&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="n"&gt;available&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;reclaimable&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Three things can be read from this.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What is subtracted is &lt;code&gt;min(half, watermark_low)&lt;/code&gt;, so &lt;strong&gt;at least half of the page cache and the reclaimable slab is counted as free&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;From the free pages, &lt;code&gt;totalreserve_pages&lt;/code&gt; is subtracted, not &lt;code&gt;watermark_low&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;The page cache term is the file LRU (&lt;code&gt;NR_ACTIVE_FILE + NR_INACTIVE_FILE&lt;/code&gt;), so &lt;strong&gt;&lt;code&gt;shmem&lt;/code&gt;, which sits on the anon LRU, is never counted as free&lt;/strong&gt;. On nodes that use tmpfs heavily, &lt;code&gt;MemAvailable&lt;/code&gt; also comes out small&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important point is that &lt;strong&gt;the two values always diverge, but which one is larger depends on the environment&lt;/strong&gt;. There are four factors behind the divergence, and they push in opposite directions.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Factor&lt;/th&gt;
&lt;th&gt;Which utilization comes out higher&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;The working set counts &lt;code&gt;active_file&lt;/code&gt; entirely as "in use"&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;The kubelet's utilization is higher&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The working set removes &lt;code&gt;inactive_file&lt;/code&gt; entirely as "free"&lt;/td&gt;
&lt;td&gt;node-exporter's utilization is higher&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;MemAvailable&lt;/code&gt; counts only about half of the page cache as "free" (&lt;code&gt;shmem&lt;/code&gt; is never counted)&lt;/td&gt;
&lt;td&gt;node-exporter's utilization is higher&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The root's usage is computed as &lt;code&gt;anon + file&lt;/code&gt;, so it includes no kernel memory such as slab or pagetables&lt;/td&gt;
&lt;td&gt;node-exporter's utilization is higher&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Three of the four factors push node-exporter's utilization higher, and only &lt;code&gt;active_file&lt;/code&gt; pushes the kubelet's utilization higher.&lt;/strong&gt; So the kubelet's utilization comes out higher only when &lt;code&gt;active_file&lt;/code&gt; has grown to several GiB and outweighs the other three; otherwise node-exporter's utilization is higher.&lt;/p&gt;

&lt;p&gt;The kubelet side is higher on nodes running workloads where &lt;code&gt;active_file&lt;/code&gt; builds up easily, for example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Batch jobs and ETL that read and write large files&lt;/li&gt;
&lt;li&gt;Services that emit a lot of logs (stdout is written to files on the node by the container runtime, so it becomes page cache)&lt;/li&gt;
&lt;li&gt;Middleware designed around the page cache (Kafka, Elasticsearch, etc.)&lt;/li&gt;
&lt;li&gt;Log collection agents that keep reading log files on the node&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What matters is the amount of file access, not the language or runtime. For example, a JVM heap is anonymous pages, so it raises RSS, not &lt;code&gt;active_file&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;On such a node it can happen that utilization computed from &lt;code&gt;MemAvailable&lt;/code&gt; looks like 33 % while the kubelet's side has reached 53 %. &lt;strong&gt;If you look at the &lt;code&gt;MemAvailable&lt;/code&gt;-based utilization and conclude there is plenty of room, the kubelet may in fact be approaching its threshold and evictions will occur.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;On the other hand, on a node where page cache has not built up much:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;MemTotal          3.820 GiB
MemAvailable      2.692 GiB   →  1 - 2.692/3.820 &lt;span class="o"&gt;=&lt;/span&gt; 29.6 %  &lt;span class="o"&gt;(&lt;/span&gt;node-exporter side&lt;span class="o"&gt;)&lt;/span&gt;
usage&lt;span class="o"&gt;(&lt;/span&gt;/&lt;span class="o"&gt;)&lt;/span&gt;          2.557 GiB   &lt;span class="o"&gt;(=&lt;/span&gt; anon + file of the root&lt;span class="o"&gt;)&lt;/span&gt;
inactive_file&lt;span class="o"&gt;(&lt;/span&gt;/&lt;span class="o"&gt;)&lt;/span&gt;  1.622 GiB
active_file&lt;span class="o"&gt;(&lt;/span&gt;/&lt;span class="o"&gt;)&lt;/span&gt;    0.227 GiB
Working Set&lt;span class="o"&gt;(&lt;/span&gt;/&lt;span class="o"&gt;)&lt;/span&gt;    0.935 GiB   →  0.935/3.820     &lt;span class="o"&gt;=&lt;/span&gt; 24.5 %  &lt;span class="o"&gt;(&lt;/span&gt;kubelet side&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;code&gt;active_file&lt;/code&gt; is only 0.23 GiB while &lt;code&gt;inactive_file&lt;/code&gt; is 1.6 GiB, so the working set removes it all as "free", and the kubelet side comes out about 5 points lower. &lt;strong&gt;The kubelet side is not always higher.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Either way, the conclusion is that &lt;strong&gt;you cannot judge whether there is headroom from only one of the values&lt;/strong&gt;. &lt;code&gt;active_file&lt;/code&gt; is an auxiliary indicator for confirming the cause of the divergence, and its value is not itself the difference. When investigating node memory pressure, also check the kubelet-view utilization with the following PromQL.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Node memory utilization from the kubelet's view (the value used in eviction decisions)&lt;/span&gt;
max by &lt;span class="o"&gt;(&lt;/span&gt;node&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  container_memory_working_set_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;*&lt;/span&gt; on&lt;span class="o"&gt;(&lt;/span&gt;instance&lt;span class="o"&gt;)&lt;/span&gt; group_left&lt;span class="o"&gt;(&lt;/span&gt;node&lt;span class="o"&gt;)&lt;/span&gt; kubelet_node_name
&lt;span class="o"&gt;)&lt;/span&gt;
/ max by &lt;span class="o"&gt;(&lt;/span&gt;node&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;kube_node_status_capacity&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;resource&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"memory"&lt;/span&gt;&lt;span class="o"&gt;})&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Amount of active page cache, the main cause of the divergence&lt;/span&gt;
max by &lt;span class="o"&gt;(&lt;/span&gt;node&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  container_memory_total_active_file_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;*&lt;/span&gt; on&lt;span class="o"&gt;(&lt;/span&gt;instance&lt;span class="o"&gt;)&lt;/span&gt; group_left&lt;span class="o"&gt;(&lt;/span&gt;node&lt;span class="o"&gt;)&lt;/span&gt; kubelet_node_name
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; The label alignment is there because &lt;strong&gt;no cAdvisor &lt;code&gt;container_*&lt;/code&gt; metric carries any label identifying the node&lt;/strong&gt;. Looking directly at the kubelet's &lt;code&gt;/metrics/cadvisor&lt;/code&gt;, the root cgroup series looks like this:&lt;/p&gt;


&lt;pre class="highlight shell"&gt;&lt;code&gt;container_memory_working_set_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;container&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;,id&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/"&lt;/span&gt;,image&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;,name&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;,namespace&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;,pod&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt; 8.92522496e+08
&lt;/code&gt;&lt;/pre&gt;


&lt;p&gt;The only thing that tells you the node is the &lt;code&gt;instance&lt;/code&gt; label Prometheus adds at scrape time, and labels such as &lt;code&gt;kubernetes_io_hostname&lt;/code&gt; are added by your Prometheus relabel configuration. &lt;strong&gt;Because the name depends on the relabel configuration (or may not exist at all), referring to it with &lt;code&gt;label_replace&lt;/code&gt; is environment specific.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;So the queries above use &lt;code&gt;kubelet_node_name&lt;/code&gt;, which the kubelet exposes on &lt;code&gt;/metrics&lt;/code&gt;.&lt;/p&gt;


&lt;pre class="highlight shell"&gt;&lt;code&gt;kubelet_node_name&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;node&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"memqos-control-plane"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt; 1
&lt;/code&gt;&lt;/pre&gt;


&lt;p&gt;This is a gauge whose value is always &lt;code&gt;1&lt;/code&gt;, and &lt;strong&gt;the node name is in the &lt;code&gt;node&lt;/code&gt; label&lt;/strong&gt;. The value itself is meaningless; it is the &lt;strong&gt;info metric&lt;/strong&gt; pattern, which carries the information in labels. The kubelet's help text says the same.&lt;/p&gt;


&lt;pre class="highlight go"&gt;&lt;code&gt;  &lt;span class="c"&gt;// NodeName is a Gauge that tracks the node's name. The count is always 1.&lt;/span&gt;
  &lt;span class="n"&gt;NodeName&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;metrics&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;NewGaugeVec&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;metrics&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GaugeOpts&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
          &lt;span class="n"&gt;Subsystem&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="n"&gt;KubeletSubsystem&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="n"&gt;Name&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;      &lt;span class="n"&gt;NodeNameKey&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="n"&gt;Help&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;      &lt;span class="s"&gt;"The node's name. The count is always 1."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;


&lt;p&gt;In OpenMetrics, &lt;code&gt;Info&lt;/code&gt; is defined as an independent metric type, with a &lt;strong&gt;mandatory &lt;code&gt;_info&lt;/code&gt; suffix and a sample value that is always &lt;code&gt;1&lt;/code&gt;&lt;/strong&gt; (&lt;a href="https://prometheus.io/docs/specs/om/open_metrics_spec/#suffixes" rel="noopener noreferrer"&gt;OpenMetrics spec - Suffixes&lt;/a&gt;).&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Info metrics are used to expose textual information which SHOULD NOT change during process lifetime.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;code&gt;kubelet_node_name&lt;/code&gt; has no &lt;code&gt;_info&lt;/code&gt; suffix and is exposed as a gauge, so &lt;strong&gt;strictly speaking it is not an OpenMetrics &lt;code&gt;Info&lt;/code&gt; type&lt;/strong&gt;. But "value always 1, information in labels" is the same usage, so this article treats it as an info-metric pattern. Similar examples are kube-state-metrics' &lt;code&gt;kube_node_info&lt;/code&gt; and node-exporter's &lt;code&gt;node_uname_info&lt;/code&gt;. Both endpoints are scraped from the same target, so &lt;code&gt;instance&lt;/code&gt; matches, and &lt;code&gt;* on(instance) group_left(node)&lt;/code&gt; brings in the &lt;code&gt;node&lt;/code&gt; label. And the resulting &lt;code&gt;node&lt;/code&gt; label has the same name as kube-state-metrics' &lt;code&gt;node&lt;/code&gt; label, so they join directly.&lt;/p&gt;

&lt;p&gt;If your environment adds its own node label by relabeling, you can use &lt;code&gt;label_replace&lt;/code&gt; as follows (replace &lt;code&gt;kubernetes_io_hostname&lt;/code&gt; with the value in your environment).&lt;/p&gt;


&lt;pre class="highlight shell"&gt;&lt;code&gt;max by &lt;span class="o"&gt;(&lt;/span&gt;node&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  label_replace&lt;span class="o"&gt;(&lt;/span&gt;
    container_memory_working_set_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;,
    &lt;span class="s2"&gt;"node"&lt;/span&gt;, &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;, &lt;span class="s2"&gt;"kubernetes_io_hostname"&lt;/span&gt;, &lt;span class="s2"&gt;"(.*)"&lt;/span&gt;
  &lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;


&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Eviction candidates are not limited to Pods in particular namespaces. When node memory gets tight, Pods in every namespace on that node are candidates. Which Pod is chosen is decided in this order: (1) whether usage exceeds the memory request, (2) Pod priority, (3) how far above the request it is. See &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/node-pressure-eviction/#pod-selection-for-kubelet-eviction" rel="noopener noreferrer"&gt;Pod selection for kubelet eviction&lt;/a&gt; for details.&lt;/p&gt;

&lt;p&gt;However, &lt;strong&gt;critical pods are excluded from selection and are never evicted at all&lt;/strong&gt; (&lt;code&gt;evictPod()&lt;/code&gt; in &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/eviction/eviction_manager.go#L633-L640" rel="noopener noreferrer"&gt;eviction_manager.go&lt;/a&gt;). "Critical pod" is a term in the official Kubernetes documentation and refers to a Pod whose PriorityClass is &lt;code&gt;system-cluster-critical&lt;/code&gt; or &lt;code&gt;system-node-critical&lt;/code&gt; (&lt;a href="https://kubernetes.io/docs/tasks/administer-cluster/guaranteed-scheduling-critical-addon-pods/" rel="noopener noreferrer"&gt;Guaranteed Scheduling For Critical Add-On Pods&lt;/a&gt; says "To mark a Pod as critical, set priorityClassName for that Pod to &lt;code&gt;system-cluster-critical&lt;/code&gt; or &lt;code&gt;system-node-critical&lt;/code&gt;.").&lt;/p&gt;


&lt;pre class="highlight go"&gt;&lt;code&gt;  &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;kubelettypes&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;IsCriticalPod&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pod&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="n"&gt;logger&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="no"&gt;nil&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"Eviction manager: cannot evict a critical pod"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"pod"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;klog&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;KObj&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pod&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="no"&gt;false&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;


&lt;p&gt;The kubelet's check is broader than this definition: &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/types/pod_update.go#L157-L169" rel="noopener noreferrer"&gt;&lt;code&gt;IsCriticalPod()&lt;/code&gt;&lt;/a&gt; is also true for static Pods and mirror Pods. The remaining condition is that &lt;code&gt;pod.Spec.Priority&lt;/code&gt; is at least &lt;code&gt;SystemCriticalPriority&lt;/code&gt; (= &lt;code&gt;2 × 1000000000&lt;/code&gt;), and since the upper limit of user-definable priority is &lt;code&gt;HighestUserDefinablePriority&lt;/code&gt; (= &lt;code&gt;1000000000&lt;/code&gt;), the only PriorityClasses that match are the two above. &lt;strong&gt;Pods with these two are not evicted, but in exchange they may not be protected in a global OOM.&lt;/strong&gt; See Which process gets killed for details.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;
  
  
  Investigating OOM and eviction
&lt;/h2&gt;

&lt;p&gt;This section covers which values to check and in what order when &lt;code&gt;OOMKilled&lt;/code&gt; or an eviction actually occurs.&lt;/p&gt;
&lt;h3&gt;
  
  
  The two OOM kill paths
&lt;/h3&gt;

&lt;p&gt;There are two paths by which a container's process gets OOM killed. Both are SIGKILLs from the kernel's OOM killer, but the conditions under which they occur and the scope they affect differ. As described earlier, memcg OOM enforces a configured limit and global OOM protects the node, so their purposes differ at the root.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Path&lt;/th&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;What gets killed&lt;/th&gt;
&lt;th&gt;Node event&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;memcg OOM&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;memory.max&lt;/code&gt; exceeded for a container, Pod or &lt;code&gt;/kubepods.slice&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Processes inside the cgroup that hit its limit&lt;/td&gt;
&lt;td&gt;Not recorded&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;global OOM&lt;/td&gt;
&lt;td&gt;Memory exhaustion of the whole node&lt;/td&gt;
&lt;td&gt;Chosen from every process on the node by &lt;code&gt;oom_score&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;SystemOOM&lt;/code&gt; is recorded if the kubelet detects it&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Even if your container is under its limit, it can be OOM killed because of another Pod on the same node.&lt;/strong&gt; That is global OOM.&lt;/p&gt;

&lt;p&gt;A Node event is recorded only by the kubelet. If &lt;code&gt;SystemOOM&lt;/code&gt; exists it confirms global OOM, but &lt;strong&gt;events are kept for only one hour by default&lt;/strong&gt;, so they become unfindable over time. So &lt;strong&gt;not finding a &lt;code&gt;SystemOOM&lt;/code&gt; does not mean it was not a global OOM&lt;/strong&gt;.&lt;/p&gt;
&lt;h4&gt;
  
  
  OOM kill from exceeding a cgroup limit (memcg OOM)
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;resources.limits.memory&lt;/code&gt; is set as the container cgroup's &lt;code&gt;memory.max&lt;/code&gt;. When a cgroup's &lt;code&gt;memory.current&lt;/code&gt; reaches &lt;code&gt;memory.max&lt;/code&gt;, this is what happens. As described above, limits are set not only on containers but also on Pods and &lt;code&gt;/kubepods.slice&lt;/code&gt;, and the same flow applies whichever level reaches its limit.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The kernel tries to reclaim memory inside that cgroup. Page cache is basically reclaimed here.&lt;/li&gt;
&lt;li&gt;If usage still does not fit under &lt;code&gt;memory.max&lt;/code&gt; after reclaim, the cgroup's OOM killer SIGKILLs processes in that cgroup.&lt;/li&gt;
&lt;li&gt;If the container's main process was killed, &lt;code&gt;lastState.terminated.reason&lt;/code&gt; becomes &lt;code&gt;OOMKilled&lt;/code&gt; and the container is recreated according to &lt;code&gt;restartPolicy&lt;/code&gt;. If only a child process was killed, the container keeps running, so neither &lt;code&gt;OOMKilled&lt;/code&gt; nor a restart is observed.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The implementation is &lt;code&gt;try_charge_memcg()&lt;/code&gt; in &lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/memcontrol.c#L2157-L2158" rel="noopener noreferrer"&gt;mm/memcontrol.c&lt;/a&gt;. It is a goto loop returning to the &lt;code&gt;retry:&lt;/code&gt; label, and proceeds in this order.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;Location&lt;/th&gt;
&lt;th&gt;What happens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1. Try to charge&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/memcontrol.c#L2177-L2178" rel="noopener noreferrer"&gt;L2177-L2178&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;If &lt;code&gt;page_counter_try_charge()&lt;/code&gt; succeeds, the allocation completes and it returns&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2. Reclaim&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/memcontrol.c#L2211-L2216" rel="noopener noreferrer"&gt;L2211-L2216&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;On failure it reclaims with &lt;code&gt;try_to_free_mem_cgroup_pages()&lt;/code&gt;, and if &lt;code&gt;mem_cgroup_margin()&lt;/code&gt; satisfies the request, goes back to 1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3. Count the retry&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/memcontrol.c#L2244-L2245" rel="noopener noreferrer"&gt;L2244-L2245&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;If reclaim is not enough, it decrements &lt;code&gt;nr_retries&lt;/code&gt; and goes back to 1. The initial value is &lt;code&gt;MAX_RECLAIM_RETRIES&lt;/code&gt; (= &lt;code&gt;16&lt;/code&gt;; &lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/internal.h#L468" rel="noopener noreferrer"&gt;mm/internal.h&lt;/a&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4. Invoke the OOM killer&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/memcontrol.c#L2259-L2264" rel="noopener noreferrer"&gt;L2259-L2264&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Only after the retries are used up does it call &lt;code&gt;mem_cgroup_oom()&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;An OOM happens only after up to 16 reclaim-and-retry rounds.&lt;/strong&gt; The check in step 2 looks like this in code.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;    &lt;span class="n"&gt;nr_reclaimed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;try_to_free_mem_cgroup_pages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;mem_over_limit&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;nr_pages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                            &lt;span class="n"&gt;gfp_mask&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;reclaim_options&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="n"&gt;psi_memstall_leave&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;pflags&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;mem_cgroup_margin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;mem_over_limit&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;nr_pages&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;goto&lt;/span&gt; &lt;span class="n"&gt;retry&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The kernel documentation for &lt;code&gt;memory.max&lt;/code&gt; says the same.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Memory usage hard limit. ... If a cgroup's memory usage reaches this limit &lt;strong&gt;and can't be reduced&lt;/strong&gt;, the OOM killer is invoked in the cgroup.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Warning: &lt;code&gt;memory.current&lt;/code&gt; reaching &lt;code&gt;memory.max&lt;/code&gt; does not by itself cause an OOM.&lt;/strong&gt; A cgroup sitting at its limit with page cache is a normal state. For the same reason, &lt;strong&gt;the working set reaching &lt;code&gt;memory.max&lt;/code&gt; is not an OOM condition either&lt;/strong&gt;. Only &lt;code&gt;inactive_file&lt;/code&gt; is subtracted from the working set, and reclaimable memory remains in the rest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;active_file&lt;/code&gt;: demoted to inactive under pressure and then reclaimed&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;slab_reclaimable&lt;/code&gt; (dentry / inode caches): in memcg reclaim too, &lt;code&gt;shrink_slab()&lt;/code&gt; is called after &lt;code&gt;shrink_lruvec()&lt;/code&gt; (&lt;code&gt;shrink_node_memcgs()&lt;/code&gt; in &lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/vmscan.c#L5917-L5920" rel="noopener noreferrer"&gt;mm/vmscan.c&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So there is no threshold of "OOM once either value reaches the limit". The actual condition is &lt;strong&gt;when the charge fails and reclaim cannot make room&lt;/strong&gt;. What cannot be reclaimed is mainly anonymous pages (without swap) and &lt;code&gt;shmem&lt;/code&gt;, so that is where to look when narrowing down.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That said, &lt;strong&gt;what the kernel compares against &lt;code&gt;memory.max&lt;/code&gt; in a memcg OOM is the whole of &lt;code&gt;memory.current&lt;/code&gt;&lt;/strong&gt;, and it is not decided by any single metric. &lt;code&gt;memory.current&lt;/code&gt; includes, besides anonymous pages, kernel memory, &lt;code&gt;shmem&lt;/code&gt; (tmpfs and shared memory), and page cache that cannot be reclaimed, or cannot be reclaimed right away.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;container_memory_rss&lt;/code&gt; is an indicator for determining whether anonymous pages are the main cause. If RSS is close to the limit you can conclude anonymous pages are the main cause, but &lt;strong&gt;RSS being below the limit is no evidence that an OOM will not occur&lt;/strong&gt;. When the gap between RSS and the working set is large, check &lt;code&gt;container_memory_usage_bytes&lt;/code&gt; (the value of &lt;code&gt;memory.current&lt;/code&gt;) together with the breakdown in &lt;code&gt;memory.stat&lt;/code&gt; on the node: &lt;code&gt;anon&lt;/code&gt;, &lt;code&gt;file&lt;/code&gt;, &lt;code&gt;shmem&lt;/code&gt;, &lt;code&gt;slab&lt;/code&gt; and so on.&lt;/p&gt;
&lt;h4&gt;
  
  
  OOM kill from node memory exhaustion (global OOM)
&lt;/h4&gt;

&lt;p&gt;When the memory of the whole node is exhausted, the kernel's OOM killer evaluates every process on the node, picks the one with the highest &lt;code&gt;oom_score&lt;/code&gt;, and SIGKILLs it. Whether an individual container exceeds its limit is irrelevant.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;However, what acts first against node memory pressure is the kubelet's eviction.&lt;/strong&gt; The kubelet evicts Pods to protect the node, and its selection puts "Pods whose usage exceeds their memory request" first (&lt;code&gt;rankMemoryPressure()&lt;/code&gt; in &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/eviction/helpers.go#L816-L821" rel="noopener noreferrer"&gt;helpers.go&lt;/a&gt;).&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;&lt;span class="c"&gt;// rankMemoryPressure orders the input pods for eviction in response to memory pressure.&lt;/span&gt;
&lt;span class="c"&gt;// It ranks by whether or not the pod's usage exceeds its requests, then by priority, and&lt;/span&gt;
&lt;span class="c"&gt;// finally by memory usage above requests.&lt;/span&gt;
&lt;span class="k"&gt;func&lt;/span&gt; &lt;span class="n"&gt;rankMemoryPressure&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pods&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;v1&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Pod&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;stats&lt;/span&gt; &lt;span class="n"&gt;statsFunc&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="n"&gt;orderedBy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;exceedMemoryRequests&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;priority&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;stats&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Sort&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pods&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;A Pod with neither &lt;code&gt;limits&lt;/code&gt; nor &lt;code&gt;requests&lt;/code&gt; always falls on the "exceeds the request" side, so it is chosen first. So if this kind of Pod makes the node tight and eviction is in time, &lt;strong&gt;it is observed as &lt;code&gt;Evicted&lt;/code&gt;, not &lt;code&gt;OOMKilled&lt;/code&gt;&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Global OOM happens first &lt;strong&gt;when eviction cannot reclaim memory as fast as memory grows&lt;/strong&gt;. There are several reasons it cannot keep up.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Factor&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;th&gt;Effect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Evaluation interval&lt;/td&gt;
&lt;td&gt;10 seconds (&lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/kubelet.go#L197" rel="noopener noreferrer"&gt;&lt;code&gt;evictionMonitoringPeriod&lt;/code&gt;&lt;/a&gt;)&lt;/td&gt;
&lt;td&gt;Detection can be up to 10 seconds late&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pods evicted per evaluation&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Only one&lt;/strong&gt; ("we kill at most a single pod during each eviction interval" in &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/eviction/eviction_manager.go#L428-L451" rel="noopener noreferrer"&gt;eviction_manager.go&lt;/a&gt;)&lt;/td&gt;
&lt;td&gt;Reclaim speed is limited to "one Pod per 10 seconds"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Waiting for cleanup&lt;/td&gt;
&lt;td&gt;Up to 30 seconds (&lt;code&gt;podCleanupTimeout&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;The next evaluation may be delayed further&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Default threshold&lt;/td&gt;
&lt;td&gt;&lt;code&gt;memory.available&amp;lt;100Mi&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Extremely thin on large nodes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;How &lt;code&gt;memory.available&lt;/code&gt; is computed&lt;/td&gt;
&lt;td&gt;From the root cgroup's &lt;code&gt;anon + file&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Contains no kernel memory, so it overestimates free memory
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The official Kubernetes documentation also says (&lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/node-pressure-eviction/#node-out-of-memory-behavior" rel="noopener noreferrer"&gt;Node out of memory behavior&lt;/a&gt;):&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If the node experiences an out of memory (OOM) event prior to the kubelet being able to reclaim memory, the node depends on the oom_killer to respond.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It is more likely when applications that allocate a lot of memory right at startup, or in response to a sudden surge in requests, share a node.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Warning: Eviction protects the node; it does not protect individual Pods from OOM.&lt;/strong&gt; Containers without &lt;code&gt;limits&lt;/code&gt; have no &lt;code&gt;memory.max&lt;/code&gt; backstop at all, so eviction is merely the last line of defense.&lt;/p&gt;

&lt;p&gt;Moreover, &lt;code&gt;rankMemoryPressure()&lt;/code&gt; looks at &lt;strong&gt;Pod priority before usage size&lt;/strong&gt;. If the offending Pod has high priority, low-priority Pods with small usage get evicted first, and global OOM can occur without the pressure being resolved. The proper countermeasure is to set &lt;code&gt;limits&lt;/code&gt; and &lt;code&gt;requests&lt;/code&gt; rather than relying on eviction.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h5&gt;
  
  
  Which process gets killed
&lt;/h5&gt;

&lt;p&gt;The kernel kills the process with the highest &lt;code&gt;oom_score&lt;/code&gt;. The kubelet sets &lt;code&gt;oom_score_adj&lt;/code&gt; (a bonus added to &lt;code&gt;oom_score&lt;/code&gt;) according to the container's QoS class, so &lt;strong&gt;the QoS class and the memory request setting directly become how likely a container is to be killed&lt;/strong&gt;. You can check the QoS class with &lt;code&gt;kubectl get pod &amp;lt;Pod name&amp;gt; -o jsonpath='{.status.qosClass}'&lt;/code&gt;, and how it is determined is described in &lt;a href="https://kubernetes.io/docs/concepts/workloads/pods/pod-qos/" rel="noopener noreferrer"&gt;Pod Quality of Service Classes&lt;/a&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Condition&lt;/th&gt;
&lt;th&gt;&lt;code&gt;oom_score_adj&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;Likelihood of being killed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;&lt;code&gt;system-node-critical&lt;/code&gt;&lt;/strong&gt; (evaluated before the QoS class)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;-997&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Least likely&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Guaranteed&lt;/td&gt;
&lt;td&gt;&lt;code&gt;-997&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Least likely&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Burstable&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;1000 - 1000 × memory request / node memory capacity&lt;/code&gt; (clamped to 3–999)&lt;/td&gt;
&lt;td&gt;The smaller the request, the more likely&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BestEffort&lt;/td&gt;
&lt;td&gt;&lt;code&gt;1000&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Most likely&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The implementation is &lt;code&gt;GetContainerOOMScoreAdjust()&lt;/code&gt; in Kubernetes' &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/qos/policy.go#L47-L59" rel="noopener noreferrer"&gt;policy.go&lt;/a&gt;. &lt;code&gt;system-node-critical&lt;/code&gt; is evaluated before the QoS class decision.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;types&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;IsNodeCriticalPod&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;pod&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="c"&gt;// Only node critical pod should be the last to get killed.&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;guaranteedOOMScoreAdj&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The Burstable clamp is implemented with a lower bound of &lt;code&gt;1000 + guaranteedOOMScoreAdj&lt;/code&gt; (= 3) and an upper bound of &lt;code&gt;besteffortOOMScoreAdj - 1&lt;/code&gt; (= 999). A Pod whose request is set smaller than its actual usage is more likely to be targeted in both eviction and global OOM.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Warning: PriorityClass works differently for eviction and for global OOM.&lt;/strong&gt; It is a misreading to think that raising priority protects you from both.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;PriorityClass&lt;/th&gt;
&lt;th&gt;Eviction&lt;/th&gt;
&lt;th&gt;global OOM (&lt;code&gt;oom_score_adj&lt;/code&gt;)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;User defined (up to &lt;code&gt;1000000000&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;The higher the value, the later it is evicted&lt;/td&gt;
&lt;td&gt;Determined by QoS class and request&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;system-cluster-critical&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Excluded (never evicted)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Determined by QoS class and request (&lt;strong&gt;not protected&lt;/strong&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;system-node-critical&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Excluded (never evicted)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Fixed at &lt;code&gt;-997&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;static / mirror Pod&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Excluded (never evicted)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Determined by QoS class and request&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The three "excluded" rows in the Eviction column are the critical pods described earlier (Pods for which &lt;code&gt;IsCriticalPod()&lt;/code&gt; is true).&lt;/p&gt;

&lt;p&gt;What &lt;code&gt;oom_score_adj&lt;/code&gt; protects is only &lt;code&gt;system-node-critical&lt;/code&gt;, for which &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/types/pod_update.go#L190-L193" rel="noopener noreferrer"&gt;&lt;code&gt;IsNodeCriticalPod()&lt;/code&gt;&lt;/a&gt; is true. &lt;strong&gt;If a &lt;code&gt;system-cluster-critical&lt;/code&gt; Pod is BestEffort, its &lt;code&gt;oom_score_adj&lt;/code&gt; is &lt;code&gt;1000&lt;/code&gt;&lt;/strong&gt;, and it is killed first in a global OOM. "I raised the priority but it was still &lt;code&gt;OOMKilled&lt;/code&gt;" comes from this path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Warning:&lt;/strong&gt; Even in a global OOM, the Pod is not deleted as in eviction. What happens next depends on &lt;code&gt;restartPolicy&lt;/code&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;Always&lt;/code&gt; or &lt;code&gt;OnFailure&lt;/code&gt;: the container is recreated on the same node. &lt;strong&gt;The Pod stays on the same node&lt;/strong&gt;, so if the cause is not resolved, OOM kills will happen again&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;Never&lt;/code&gt;: the container is not recreated and the Pod remains &lt;code&gt;Failed&lt;/code&gt;. If a controller such as a Job creates a successor Pod, it may be placed on a different node&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If &lt;code&gt;OOMKilled&lt;/code&gt; repeats even though there is headroom under the limit, suspect this path.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Because the OOM kill frees memory, the pressure may already be gone when the kubelet next evaluates free memory. In that case the node gets no &lt;code&gt;MemoryPressure&lt;/code&gt; condition and no &lt;code&gt;node.kubernetes.io/memory-pressure:NoSchedule&lt;/code&gt; taint, and scheduling of new Pods does not stop. Conversely, if the pressure continues, a later evaluation sets the condition and taint, and eviction also occurs. Note that it is decided by &lt;code&gt;memory.available&lt;/code&gt; at evaluation time, not by whether an OOM kill happened.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3&gt;
  
  
  Telling which path occurred
&lt;/h3&gt;
&lt;h4&gt;
  
  
  Using Kubernetes information
&lt;/h4&gt;

&lt;p&gt;Start by narrowing down with what you can check using &lt;code&gt;kubectl&lt;/code&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Where to look&lt;/th&gt;
&lt;th&gt;What it tells you&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;Last State&lt;/code&gt; in &lt;code&gt;kubectl describe pod&lt;/code&gt;, or &lt;code&gt;lastState.terminated&lt;/code&gt; in &lt;code&gt;kubectl get pod -o yaml&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;That an OOM kill happened (&lt;code&gt;reason: OOMKilled&lt;/code&gt;). &lt;strong&gt;It cannot distinguish the path&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Events in &lt;code&gt;kubectl describe node &amp;lt;node name&amp;gt;&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;If &lt;code&gt;SystemOOM&lt;/code&gt; is recorded, it is a global OOM.&lt;/strong&gt; If there is no record, that does not make it a memcg OOM; it just means you cannot tell&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;There is no OOM-specific event on the Pod side. You only see things like &lt;code&gt;BackOff&lt;/code&gt; accompanying the restart.&lt;/p&gt;

&lt;p&gt;Check it as follows. &lt;code&gt;SystemOOM&lt;/code&gt; is an event attached to the Node, and since Node is not a namespaced object, the event itself is recorded in the &lt;code&gt;default&lt;/code&gt; namespace.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 1. Find which node the OOMKilled Pod is on&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl get pod &lt;span class="nt"&gt;-n&lt;/span&gt; &amp;lt;namespace&amp;gt; &amp;lt;Pod name&amp;gt; &lt;span class="nt"&gt;-o&lt;/span&gt; wide
NAME       READY   STATUS    RESTARTS      AGE   IP            NODE
myapp-...  1/1     Running   5 &lt;span class="o"&gt;(&lt;/span&gt;2m ago&lt;span class="o"&gt;)&lt;/span&gt;    1h    10.0.0.1      node-1

&lt;span class="c"&gt;# 2. Check whether that node's events contain SystemOOM&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl describe node node-1
...
Events:
  Type     Reason     Age   From     Message
  &lt;span class="nt"&gt;----&lt;/span&gt;     &lt;span class="nt"&gt;------&lt;/span&gt;     &lt;span class="nt"&gt;----&lt;/span&gt;  &lt;span class="nt"&gt;----&lt;/span&gt;     &lt;span class="nt"&gt;-------&lt;/span&gt;
  Warning  SystemOOM  3m    kubelet  System OOM encountered, victim process: java, pid: 12345

&lt;span class="c"&gt;# To list only events&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl get events &lt;span class="nt"&gt;-n&lt;/span&gt; default &lt;span class="nt"&gt;--field-selector&lt;/span&gt; &lt;span class="nv"&gt;reason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;SystemOOM,involvedObject.name&lt;span class="o"&gt;=&lt;/span&gt;node-1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;If &lt;code&gt;SystemOOM&lt;/code&gt; exists, a global OOM definitely happened on that node. However, the victim is not necessarily the container you are investigating, so also check that the time of &lt;code&gt;SystemOOM&lt;/code&gt; corresponds to the container's termination time (&lt;code&gt;lastState.terminated.finishedAt&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not finding &lt;code&gt;SystemOOM&lt;/code&gt; is not evidence of memcg OOM.&lt;/strong&gt; The same state results from event TTL expiry, a kubelet restart, or a missed event.&lt;/p&gt;

&lt;p&gt;The event retention period is set by kube-apiserver's &lt;code&gt;--event-ttl&lt;/code&gt;, and &lt;strong&gt;the default is 1 hour&lt;/strong&gt; (&lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/controlplane/apiserver/options/options.go#L129" rel="noopener noreferrer"&gt;options.go&lt;/a&gt;). So &lt;strong&gt;a &lt;code&gt;SystemOOM&lt;/code&gt; older than an hour cannot be seen with &lt;code&gt;kubectl&lt;/code&gt;&lt;/strong&gt;. If you forward events to an external logging platform, check there.&lt;/p&gt;

&lt;p&gt;When you cannot tell, infer from the following:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Whether &lt;code&gt;container_memory_usage_bytes&lt;/code&gt; was close to the limit at that time (if so, memcg OOM is likely)&lt;/li&gt;
&lt;li&gt;Whether the kubelet-view node memory utilization was high at that time (if so, global OOM is likely)&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;
  
  
  Why OOMKilled cannot distinguish the paths
&lt;/h4&gt;

&lt;p&gt;When a task exits with code 137, the container runtime (containerd) checks &lt;code&gt;oom_kill&lt;/code&gt; in the cgroup's &lt;code&gt;memory.events&lt;/code&gt;, and if it is 1 or more, it updates the termination reason to &lt;code&gt;OOMKilled&lt;/code&gt; (&lt;a href="https://github.com/containerd/containerd/blob/v2.3.5/internal/cri/server/events.go#L194-L238" rel="noopener noreferrer"&gt;events.go&lt;/a&gt;). The check is &lt;a href="https://github.com/containerd/containerd/blob/v2.3.5/internal/cri/server/events.go#L399-L427" rel="noopener noreferrer"&gt;&lt;code&gt;oomMetricsEventOccurred()&lt;/code&gt;&lt;/a&gt; in the same file.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;    &lt;span class="k"&gt;switch&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;taskMetricsAny&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;type&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;cg1&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Metrics&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GetMemoryOomControl&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GetOomKill&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="m"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt;
    &lt;span class="k"&gt;case&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;cg2&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Metrics&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;v&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GetMemoryEvents&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;GetOomKill&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="m"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The definition of &lt;code&gt;oom_kill&lt;/code&gt; in &lt;code&gt;memory.events&lt;/code&gt; is "The number of processes belonging to this cgroup killed by &lt;strong&gt;any kind of&lt;/strong&gt; OOM killer". So a process chosen and killed by the kernel in a global OOM is counted too, and &lt;strong&gt;both paths produce &lt;code&gt;OOMKilled&lt;/code&gt;&lt;/strong&gt;.&lt;/p&gt;
&lt;h4&gt;
  
  
  Why SystemOOM is recorded only for global OOM
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;SystemOOM&lt;/code&gt; is an event the kubelet records on the Node object. The kubelet parses OOM messages in the kernel log and &lt;strong&gt;records it only when the cgroup where the OOM happened is judged to be the root (&lt;code&gt;/&lt;/code&gt;)&lt;/strong&gt; (&lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/oom/oom_watcher_linux.go#L72-L98" rel="noopener noreferrer"&gt;oom_watcher_linux.go&lt;/a&gt;).&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="k"&gt;range&lt;/span&gt; &lt;span class="n"&gt;outStream&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="c"&gt;// Count every OOM kill per container to back the&lt;/span&gt;
            &lt;span class="c"&gt;// container_oom_events_total metric.&lt;/span&gt;
            &lt;span class="n"&gt;recordOOMKill&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ContainerName&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;VictimContainerName&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;recordEventContainerName&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="o"&gt;...&lt;/span&gt;
                &lt;span class="n"&gt;ow&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;recorder&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;WithLogger&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;logger&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Eventf&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ref&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;v1&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;EventTypeWarning&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;systemOOMEvent&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s"&gt;"%s"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;eventMsg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;That this lines up exactly with global OOM comes from the combination of the kernel log format and the parser. For an OOM caused by a memcg limit, the kernel prints &lt;code&gt;oom_memcg=&amp;lt;cgroup path&amp;gt;&lt;/code&gt;, whereas for a global OOM it prints &lt;code&gt;,global_oom&lt;/code&gt; instead (&lt;code&gt;mem_cgroup_print_oom_context()&lt;/code&gt; in &lt;a href="https://github.com/torvalds/linux/blob/v6.12/mm/memcontrol.c#L1503-L1507" rel="noopener noreferrer"&gt;mm/memcontrol.c&lt;/a&gt;).&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight c"&gt;&lt;code&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memcg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="n"&gt;pr_cont&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;",oom_memcg="&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="n"&gt;pr_cont_cgroup_path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memcg&lt;/span&gt;&lt;span class="o"&gt;-&amp;gt;&lt;/span&gt;&lt;span class="n"&gt;css&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;cgroup&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt;
        &lt;span class="nf"&gt;pr_cont&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;",global_oom"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The parser side only matches the format containing &lt;code&gt;oom_memcg=&lt;/code&gt; (&lt;a href="https://github.com/google/cadvisor/blob/v0.57.0/utils/oomparser/oomparser.go#L30-L36" rel="noopener noreferrer"&gt;oomparser.go&lt;/a&gt;).&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;    &lt;span class="n"&gt;containerRegexp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;regexp&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;MustCompile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s"&gt;`oom-kill:constraint=(.*),nodemask=(.*),cpuset=(.*),mems_allowed=(.*),oom_memcg=(.*),task_memcg=(.*),task=(.*),pid=(.*),uid=(.*)`&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;So no cgroup can be extracted from a global OOM message, and the initial value &lt;code&gt;/&lt;/code&gt; of &lt;code&gt;OomInstance&lt;/code&gt; is left as is (&lt;code&gt;StreamOoms&lt;/code&gt; in &lt;a href="https://github.com/google/cadvisor/blob/v0.57.0/utils/oomparser/oomparser.go#L126-L130" rel="noopener noreferrer"&gt;oomparser.go&lt;/a&gt;).&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;            &lt;span class="n"&gt;oomCurrentInstance&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;OomInstance&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="n"&gt;ContainerName&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;       &lt;span class="s"&gt;"/"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="n"&gt;VictimContainerName&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s"&gt;"/"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="n"&gt;TimeOfDeath&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;         &lt;span class="n"&gt;msg&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Timestamp&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;In other words, the presence of &lt;code&gt;SystemOOM&lt;/code&gt; is not "the result of directly determining the path" but "the result of whether the kernel log could be parsed". Still, in practice it is enough to understand that &lt;strong&gt;if it was recorded, you can conclude it was a global OOM&lt;/strong&gt;, while &lt;strong&gt;its absence is no evidence at all&lt;/strong&gt;.&lt;/p&gt;
&lt;h4&gt;
  
  
  How it appears in metrics
&lt;/h4&gt;

&lt;p&gt;Metrics do not distinguish the path with a label. The kernel's &lt;code&gt;__oom_kill_process()&lt;/code&gt; records the same event regardless of the path.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Emitted by&lt;/th&gt;
&lt;th&gt;memcg OOM&lt;/th&gt;
&lt;th&gt;global OOM&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;node_vmstat_oom_kill&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;node-exporter (&lt;code&gt;oom_kill&lt;/code&gt; in &lt;code&gt;/proc/vmstat&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Increases&lt;/td&gt;
&lt;td&gt;Increases&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_oom_events_total&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;kubelet (result of kernel log parsing)&lt;/td&gt;
&lt;td&gt;The series of the container the killed process belonged to increases&lt;/td&gt;
&lt;td&gt;The &lt;code&gt;id="/"&lt;/code&gt; series increases, not a container series&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;oom_kill&lt;/code&gt; in cgroup v2 &lt;code&gt;memory.events&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;(a cgroup file, not a metric; containerd reads it)&lt;/td&gt;
&lt;td&gt;Increases&lt;/td&gt;
&lt;td&gt;Increases&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Only &lt;code&gt;container_oom_events_total&lt;/code&gt; behaves differently because, as described earlier, the parser cannot extract a cgroup in a global OOM and treats it as the root (&lt;code&gt;/&lt;/code&gt;). This is only a result of whether the log could be parsed, so do not use it to determine the path; decide by the presence of &lt;code&gt;SystemOOM&lt;/code&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Although &lt;code&gt;container_oom_events_total&lt;/code&gt; has &lt;code&gt;container_&lt;/code&gt; in its name, in v1.37 it is counted by the kubelet, not cAdvisor (&lt;code&gt;recordOOMKill()&lt;/code&gt; in &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/oom/oom_counter.go#L40-L44" rel="noopener noreferrer"&gt;oom_counter.go&lt;/a&gt;). The meaning of the value does not change, but you have to look elsewhere when reading the code.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4&gt;
  
  
  What to look at after narrowing down
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;If you judge it to be a global OOM&lt;/strong&gt;, the cause is memory exhaustion of the whole node. Compare the memory requests and actual usage of the Pods placed on the same node, and check whether any Pod has too small a request&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you judge it to be a memcg OOM&lt;/strong&gt;, review that container's memory limit and its actual usage. If the container seems to have headroom under its limit, it may have reached the limit of the Pod as a whole or of &lt;code&gt;/kubepods.slice&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you cannot decide&lt;/strong&gt;, keep both possibilities open and check both the container's limit and usage and the node's memory utilization&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  Investigation workflow
&lt;/h3&gt;

&lt;p&gt;When investigating memory-related events, checking in the following order makes it easier to narrow down the cause.&lt;/p&gt;
&lt;h4&gt;
  
  
  1. When a Pod was OOMKilled
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Check whether &lt;code&gt;container_memory_usage_bytes&lt;/code&gt; is close to the limit. memcg OOM is evaluated against this value&lt;/li&gt;
&lt;li&gt;Compare with &lt;code&gt;container_memory_rss&lt;/code&gt; to tell whether anonymous pages are the main cause. An OOM can occur even when RSS is low&lt;/li&gt;
&lt;li&gt;If the gap between the working set or usage and RSS is large, page cache, &lt;code&gt;shmem&lt;/code&gt; or kernel memory may be pushing usage up&lt;/li&gt;
&lt;li&gt;If &lt;code&gt;OOMKilled&lt;/code&gt; occurs despite headroom under the limit, suspect global OOM. If the node's events have a &lt;code&gt;SystemOOM&lt;/code&gt;, it is a global OOM. If there is none, it is not necessarily a memcg OOM&lt;/li&gt;
&lt;li&gt;If you see no spike in the metrics, suspect a surge shorter than the scrape interval. See My working set is below the limit, but the container was OOMKilled for what to check&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;
  
  
  2. When a Pod was evicted
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;If the message is &lt;code&gt;The node was low on resource: memory.&lt;/code&gt;, node-wide memory pressure is the cause&lt;/li&gt;
&lt;li&gt;Use the PromQL above to check the kubelet-view node memory utilization. The node-exporter-based utilization alone is not enough to judge&lt;/li&gt;
&lt;li&gt;Check whether a Pod with &lt;strong&gt;a small (or unset) memory request relative to actual usage&lt;/strong&gt; is on the same node. Compare each Pod's &lt;code&gt;container_memory_working_set_bytes&lt;/code&gt; (cAdvisor) against &lt;code&gt;kube_pod_resource_request{resource="memory"}&lt;/code&gt; (kube-scheduler). Note that Pods with large requests make the scheduler keep placement density low, so they are unlikely to be the cause of pressure&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;
  
  
  3. To prevent recurrence
&lt;/h4&gt;

&lt;ul&gt;
&lt;li&gt;Set a memory request that matches actual usage. It has the largest effect. PromQL for checking is collected in Align memory request with actual usage
&lt;/li&gt;
&lt;li&gt;For containers where memcg OOM occurs, also review the limit&lt;/li&gt;
&lt;li&gt;Control placement with Pod anti-affinity or topology spread constraints so that Pods with large memory usage do not gather on the same node&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;
  
  
  Checking with PSI (Pressure Stall Information)
&lt;/h4&gt;

&lt;p&gt;This is an auxiliary indicator for the case where things are slow although usage has not reached the limit. It shows how long processes were actually stalled waiting for memory reclaim.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  How to read PSI and PromQL
  &lt;p&gt;You can obtain the time a container was stalled waiting for memory allocation as PSI. Even when usage has not reached the limit, you can detect being made to wait by memory reclaim.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Emitted by&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_pressure_memory_waiting_seconds_total&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;Time some processes were waiting for memory (&lt;code&gt;total&lt;/code&gt; of &lt;code&gt;some&lt;/code&gt; in &lt;code&gt;memory.pressure&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_pressure_memory_stalled_seconds_total&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet)&lt;/td&gt;
&lt;td&gt;Time all processes were stalled waiting for memory (&lt;code&gt;total&lt;/code&gt; of &lt;code&gt;full&lt;/code&gt; in &lt;code&gt;memory.pressure&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; These metrics are controlled by the &lt;code&gt;KubeletPSI&lt;/code&gt; feature gate, which &lt;strong&gt;went GA in v1.36 and is enabled by default (it cannot be disabled)&lt;/strong&gt;. The kernel must support PSI, though, and the kubelet decides this by whether &lt;code&gt;cpu.pressure&lt;/code&gt; exists at the cgroup root (&lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/cadvisor/cadvisor_linux.go#L172-L187" rel="noopener noreferrer"&gt;cadvisor_linux.go&lt;/a&gt;). If the metrics do not appear, check whether the kernel was built with &lt;code&gt;CONFIG_PSI=y&lt;/code&gt; and whether &lt;code&gt;psi=0&lt;/code&gt; is set as a kernel parameter.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;memory.pressure&lt;/code&gt; also has &lt;code&gt;avg10&lt;/code&gt; / &lt;code&gt;avg60&lt;/code&gt; / &lt;code&gt;avg300&lt;/code&gt;, "the average over the last N seconds (%)", but &lt;strong&gt;only the cumulative value (&lt;code&gt;total&lt;/code&gt;) is exposed as a Prometheus metric&lt;/strong&gt;. So use &lt;code&gt;rate()&lt;/code&gt; to see a ratio.&lt;/p&gt;

&lt;p&gt;The unit is seconds, so the result of &lt;code&gt;rate()&lt;/code&gt; is "how many seconds of waiting per second", a ratio from 0 to 1. Multiply by 100 to get a percentage directly.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Per Pod, the share of time (%) all processes were stalled waiting for memory&lt;/span&gt;
100 &lt;span class="k"&gt;*&lt;/span&gt; max by &lt;span class="o"&gt;(&lt;/span&gt;namespace, pod&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  rate&lt;span class="o"&gt;(&lt;/span&gt;container_pressure_memory_stalled_seconds_total&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;container&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, pod!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="o"&gt;}[&lt;/span&gt;5m]&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Per Pod, the share of time (%) some processes were waiting for memory&lt;/span&gt;
100 &lt;span class="k"&gt;*&lt;/span&gt; max by &lt;span class="o"&gt;(&lt;/span&gt;namespace, pod&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  rate&lt;span class="o"&gt;(&lt;/span&gt;container_pressure_memory_waiting_seconds_total&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;container&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, pod!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="o"&gt;}[&lt;/span&gt;5m]&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When there is plenty of memory, neither increases, so &lt;code&gt;rate()&lt;/code&gt; stays at 0. Under pressure, &lt;strong&gt;&lt;code&gt;waiting&lt;/code&gt; (some) rises first, and if it gets worse &lt;code&gt;stalled&lt;/code&gt; (full) follows&lt;/strong&gt;. &lt;code&gt;waiting&lt;/code&gt; counts even when only some processes waited, so &lt;strong&gt;to judge severity, use &lt;code&gt;stalled&lt;/code&gt;&lt;/strong&gt;. It is the time all processes in the container were stopped, so if it shows a value, throughput is directly affected.&lt;/p&gt;

&lt;p&gt;To see memory pressure on the whole node, use the root cgroup series.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Share of time (%) the whole node was completely stalled waiting for memory&lt;/span&gt;
100 &lt;span class="k"&gt;*&lt;/span&gt; max by &lt;span class="o"&gt;(&lt;/span&gt;node&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  rate&lt;span class="o"&gt;(&lt;/span&gt;container_pressure_memory_stalled_seconds_total&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/"&lt;/span&gt;&lt;span class="o"&gt;}[&lt;/span&gt;5m]&lt;span class="o"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;*&lt;/span&gt; on&lt;span class="o"&gt;(&lt;/span&gt;instance&lt;span class="o"&gt;)&lt;/span&gt; group_left&lt;span class="o"&gt;(&lt;/span&gt;node&lt;span class="o"&gt;)&lt;/span&gt; kubelet_node_name
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Containers that are struggling although they are below their limit can be found by combining PSI with utilization. &lt;code&gt;container_spec_memory_limit_bytes&lt;/code&gt; is the container limit exposed by cAdvisor.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Containers whose usage is under 80 % of the limit but that are stalled waiting for memory&lt;/span&gt;
&lt;span class="o"&gt;(&lt;/span&gt;
  100 &lt;span class="k"&gt;*&lt;/span&gt; rate&lt;span class="o"&gt;(&lt;/span&gt;container_pressure_memory_stalled_seconds_total&lt;span class="o"&gt;{&lt;/span&gt;container!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, image!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="o"&gt;}[&lt;/span&gt;5m]&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; 1
&lt;span class="o"&gt;)&lt;/span&gt;
and
&lt;span class="o"&gt;(&lt;/span&gt;
  container_memory_usage_bytes&lt;span class="o"&gt;{&lt;/span&gt;container!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, image!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
  / container_spec_memory_limit_bytes&lt;span class="o"&gt;{&lt;/span&gt;container!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, image!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt; &amp;lt; 0.8
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;and&lt;/code&gt; keeps only series whose label sets match exactly on both sides. A container with no limit has &lt;code&gt;container_spec_memory_limit_bytes&lt;/code&gt; of 0, so the division gives &lt;code&gt;+Inf&lt;/code&gt;, which does not satisfy &lt;code&gt;&amp;lt; 0.8&lt;/code&gt;, and it is excluded automatically.&lt;/p&gt;

&lt;p&gt;Here is an example alert. It catches &lt;code&gt;stalled&lt;/code&gt; that keeps occurring.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;alert&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ContainerMemoryStalled&lt;/span&gt;
  &lt;span class="c1"&gt;# Fires when the 5-minute average of fully stalled time is 5 % or more for 10 minutes&lt;/span&gt;
  &lt;span class="na"&gt;expr&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;100 * max by (namespace, pod) (&lt;/span&gt;
      &lt;span class="s"&gt;rate(container_pressure_memory_stalled_seconds_total{container="", image="", pod!=""}[5m])&lt;/span&gt;
    &lt;span class="s"&gt;) &amp;gt; 5&lt;/span&gt;
  &lt;span class="na"&gt;for&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;10m&lt;/span&gt;
  &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;severity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;warning&lt;/span&gt;
  &lt;span class="na"&gt;annotations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;summary&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;A&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Pod&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;is&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;stalled&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;waiting&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;memory&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;allocation&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;(add&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;namespace&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;pod&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;labels&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;your&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;own&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;template)"&lt;/span&gt;
    &lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;It&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;may&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;be&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;made&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;wait&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;by&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;reclaim&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;even&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;though&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;it&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;has&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;not&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;reached&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;its&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;memory&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;limit.&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Check&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;gap&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;between&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;usage&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;RSS&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;of&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;page&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;cache."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The right threshold of &lt;code&gt;5 %&lt;/code&gt; depends on the environment. First look at the distribution of values without an alert to understand the normal level, then decide.&lt;/p&gt;

&lt;p&gt;For details on PSI see &lt;a href="https://kubernetes.io/docs/reference/instrumentation/understand-psi-metrics/" rel="noopener noreferrer"&gt;Understand Pressure Stall Information (PSI) Metrics&lt;/a&gt;.&lt;/p&gt;



&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do about it
&lt;/h2&gt;

&lt;p&gt;This is what to do once you know the cause. &lt;strong&gt;The most effective way to reduce OOM kills and evictions together is to right-size the memory request&lt;/strong&gt;, so that is the focus.&lt;/p&gt;

&lt;h3&gt;
  
  
  Align memory request with actual usage
&lt;/h3&gt;

&lt;p&gt;As we have seen, &lt;strong&gt;the most effective way to keep OOMs down is to set the memory request to match actual usage&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;resources.requests.memory&lt;/code&gt; is not set in the cgroup, so it does not directly limit how much memory a container can use. It nevertheless matters because the request determines how likely OOM kills and evictions are, in these three ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Placement (how likely both global OOM and eviction are)&lt;/strong&gt;: the scheduler decides placement by looking only at requests, not actual usage. A Pod with no request, or one smaller than actual usage, is placed while overestimating the node's free memory. As a result, Pods with large memory usage concentrate on the same node and the whole node becomes tight, &lt;strong&gt;inviting both global OOM and eviction&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Eviction candidate selection&lt;/strong&gt;: the kubelet preferentially evicts Pods whose actual usage exceeds their memory request. If you set the request smaller than actual usage, that Pod is evicted more easily&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;oom_score_adj&lt;/code&gt; (the order of kills in a global OOM)&lt;/strong&gt;: as described earlier, the Burstable &lt;code&gt;oom_score_adj&lt;/code&gt; is &lt;code&gt;1000 - 1000 × memory request / node memory capacity&lt;/code&gt;. &lt;strong&gt;The smaller the request, the larger the &lt;code&gt;oom_score_adj&lt;/code&gt;, and the more likely the Pod is chosen as a kill target in a global OOM&lt;/strong&gt;. With no request it is BestEffort and &lt;code&gt;1000&lt;/code&gt;, so it is killed first&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;"Whether the event happens" and "which Pod is chosen" are separate matters&lt;/strong&gt;, so organized by effect:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;How the request acts&lt;/th&gt;
&lt;th&gt;memcg OOM occurs&lt;/th&gt;
&lt;th&gt;global OOM occurs&lt;/th&gt;
&lt;th&gt;global OOM target selection&lt;/th&gt;
&lt;th&gt;Eviction occurs&lt;/th&gt;
&lt;th&gt;Eviction target selection&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Placement (scheduler)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;oom_score_adj&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Eviction candidate selection&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;○&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Only placement affects whether an event occurs&lt;/strong&gt;, and it affects both global OOM and eviction, because both start from node pressure. The effect is indirect, though: the request does not appear in the computation of &lt;code&gt;memory.available&lt;/code&gt; (&lt;code&gt;makeMemoryAvailableSignalObservation()&lt;/code&gt; mentioned earlier), so raising a request does not relax the threshold. The node just becomes less likely to get tight as a result of lower placement density.&lt;/p&gt;

&lt;p&gt;The other two only affect &lt;strong&gt;whether your Pod is chosen&lt;/strong&gt; after the pressure has built up. And &lt;strong&gt;the only thing the request does not affect is memcg OOM&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Conversely, just aligning the request with actual usage gives two effects at once: "the node is less likely to get tight" and "even when it does, your Pod is less likely to be chosen". Unlike increasing the limit, right-sizing the request does not change the total memory usable across the node, so it is also less likely to translate into cost.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Warning:&lt;/strong&gt; Right-sizing the request does not help with memcg OOM (exceeding a container's limit). To reduce memcg OOM you need to review the limit. This section is about &lt;strong&gt;measures to reduce global OOM and eviction&lt;/strong&gt;. Use Telling which path occurred to determine which one is happening.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h4&gt;
  
  
  PromQL for comparing requests and actual usage
&lt;/h4&gt;

&lt;p&gt;You can check whether a request matches actual usage with PromQL. I included queries for the per-Pod ratio, for finding Pods with no request, for per-node comparison, and for deriving a recommended value from past usage. &lt;strong&gt;I also explain the pitfalls that come from where the metrics originate (such as double counting).&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  PromQL for checking and notes on filters
  &lt;p&gt;&lt;strong&gt;What to know before comparing&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A Pod's request is available as &lt;code&gt;kube_pod_resource_request&lt;/code&gt;, which kube-scheduler exposes on its &lt;code&gt;/metrics/resources&lt;/code&gt; endpoint. You compare it with actual usage (&lt;code&gt;container_memory_working_set_bytes&lt;/code&gt;), but the two come from different sources, so dividing them as is gives the wrong result. Keep these five points in mind.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Point&lt;/th&gt;
&lt;th&gt;Details&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Scrape configuration is needed&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;kube_pod_resource_request&lt;/code&gt; is on kube-scheduler's secure port (default 10259) at &lt;code&gt;/metrics/resources&lt;/code&gt;. It is a different endpoint from &lt;code&gt;/metrics&lt;/code&gt;, so Prometheus needs explicit configuration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;No series for Pods with no request&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Because the collector does &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/scheduler/metrics/resources/resources.go#L162-L177" rel="noopener noreferrer"&gt;&lt;code&gt;if val.IsZero() { return }&lt;/code&gt;&lt;/a&gt;, Pods whose request is 0 or unset do not appear in the metric. &lt;strong&gt;The most problematic Pods drop out of the query&lt;/strong&gt;, so find them separately with &lt;code&gt;unless&lt;/code&gt; below&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Aggregated per Pod&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;kube_pod_resource_request&lt;/code&gt; is the total for the whole Pod and has no &lt;code&gt;container&lt;/code&gt; label. Actual usage must also be aligned to Pod level before comparing (how is shown below)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Duplicate scrapes&lt;/td&gt;
&lt;td&gt;kube-scheduler runs multiple replicas in an HA setup, so the same Pod's series may be scraped several times with different &lt;code&gt;instance&lt;/code&gt; values. Summing them double counts, so &lt;strong&gt;normalize with &lt;code&gt;max by (...)&lt;/code&gt;&lt;/strong&gt; before aggregating&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminated Pods are excluded&lt;/td&gt;
&lt;td&gt;Pods that are Succeeded / Failed are not included&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Which metric gives per-Pod memory usage&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;code&gt;kube_pod_resource_request&lt;/code&gt; is a per-Pod value, so actual usage must also be aligned to Pod level. You may wonder "doesn't kube-state-metrics have a Pod memory usage metric?", but &lt;strong&gt;it does not&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;kube-state-metrics converts Kubernetes API objects (spec and status) directly into metrics and does not handle actual resource usage at all. It provides &lt;strong&gt;configured values written in manifests&lt;/strong&gt;, such as &lt;code&gt;kube_pod_container_resource_requests&lt;/code&gt; / &lt;code&gt;kube_pod_container_resource_limits&lt;/code&gt;, but usage is out of scope. Usage requires reading cgroups, which is the job of cAdvisor and the kubelet.&lt;/p&gt;

&lt;p&gt;There are three ways to get actual per-Pod usage.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Emitted by&lt;/th&gt;
&lt;th&gt;Characteristics&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sum the containers&lt;/td&gt;
&lt;td&gt;&lt;code&gt;sum by (namespace, pod) (container_memory_working_set_bytes{...})&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (&lt;code&gt;/metrics/cadvisor&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;Works in any environment. Double counts if the filter is wrong (see below)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Look at the Pod's cgroup directly&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;container_memory_working_set_bytes{container="", image="", pod!=""}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (&lt;code&gt;/metrics/cadvisor&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;No summing needed. Slightly larger than the container total because it includes the pause container and Pod-level cgroup charges. &lt;code&gt;image=""&lt;/code&gt; is required (see below)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The kubelet's per-Pod metric&lt;/td&gt;
&lt;td&gt;&lt;code&gt;pod_memory_working_set_bytes{namespace, pod}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;kubelet &lt;code&gt;/metrics/resource&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;The most straightforward, but it is on a different path from &lt;code&gt;/metrics/cadvisor&lt;/code&gt;, so you must add a scrape configuration to Prometheus&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The second and third are the same value. The kubelet gets &lt;code&gt;pod_memory_working_set_bytes&lt;/code&gt; &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/stats/cadvisor_stats_provider.go#L156-L168" rel="noopener noreferrer"&gt;from the Pod-level cgroup&lt;/a&gt;, not as a sum of containers.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;        &lt;span class="n"&gt;podUID&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;types&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;UID&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;podStats&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;PodRef&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;UID&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="c"&gt;// Lookup the pod-level cgroup's CPU and memory stats&lt;/span&gt;
        &lt;span class="n"&gt;podInfo&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;getCadvisorPodInfoFromPodUID&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;podUID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;allInfos&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;podInfo&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="no"&gt;nil&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;cpu&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;memory&lt;/span&gt; &lt;span class="o"&gt;:=&lt;/span&gt; &lt;span class="n"&gt;cadvisorInfoToCPUandMemoryStats&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;podInfo&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;podStats&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;CPU&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cpu&lt;/span&gt;
            &lt;span class="n"&gt;podStats&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Memory&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;memory&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;pod_memory_working_set_bytes&lt;/code&gt; is a STABLE metric, so if Prometheus scrapes &lt;code&gt;/metrics/resource&lt;/code&gt; in your environment, using it is the simplest.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Actual per-Pod usage (when scraping the kubelet's /metrics/resource)&lt;/span&gt;
pod_memory_working_set_bytes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you do not scrape &lt;code&gt;/metrics/resource&lt;/code&gt;, specify the Pod cgroup series directly. &lt;code&gt;/metrics/resource&lt;/code&gt; is a different path from &lt;code&gt;/metrics/cadvisor&lt;/code&gt;, so it is not collected unless you add it to the Prometheus configuration.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Actual per-Pod usage (looking at the Pod's cgroup series directly)&lt;/span&gt;
max by &lt;span class="o"&gt;(&lt;/span&gt;namespace, pod&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  container_memory_working_set_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;container&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, pod!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;container=""&lt;/code&gt; gets the Pod's cgroup because when the kubelet labels metrics it &lt;strong&gt;also attaches &lt;code&gt;pod&lt;/code&gt; and &lt;code&gt;namespace&lt;/code&gt; to the Pod's cgroup series&lt;/strong&gt;. The code explains it as "Associate pod cgroup with pod so we have an accurate accounting of sandbox" (&lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/server/server.go#L1481-L1495" rel="noopener noreferrer"&gt;server.go&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Warning: Do not omit &lt;code&gt;image=""&lt;/code&gt;.&lt;/strong&gt; If you filter with only &lt;code&gt;container=""&lt;/code&gt;, &lt;strong&gt;the series of the pause container (sandbox) also matches&lt;/strong&gt; along with the Pod's cgroup. In containerd environments the pause container's &lt;code&gt;container&lt;/code&gt; label is an empty string and its &lt;code&gt;image&lt;/code&gt; label holds the pause image.&lt;/p&gt;

&lt;p&gt;For example, in a cluster running 100 Pods, the series count splits into these three kinds.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Selector&lt;/th&gt;
&lt;th&gt;Series&lt;/th&gt;
&lt;th&gt;What it is&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;{container="", pod!=""}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;Pod cgroups + pause containers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;{container="", image="", pod!=""}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Pod cgroups only&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;{container="", image!="", pod!=""}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;Pause containers only&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;With &lt;code&gt;max by (namespace, pod)&lt;/code&gt; the Pod's cgroup is larger, so the correct value comes back, but with &lt;code&gt;sum by&lt;/code&gt; the pause part is added and it is double counted. For one Pod it looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;Pod cgroup                666,116,096   ← correct value
Sum of real containers     665,882,624
pause container                225,280
&lt;span class="nb"&gt;sum&lt;/span&gt;&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;container&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;,pod!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;  666,341,376   ← Pod cgroup + pause, double counted
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you do &lt;code&gt;sum by (namespace, pod)&lt;/code&gt; without a &lt;code&gt;container&lt;/code&gt; filter, each container's share and the Pod cgroup's share are both added, so the value is almost doubled. When summing containers, narrow to real containers only, as follows.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Actual per-Pod usage (when summing containers)&lt;/span&gt;
&lt;span class="nb"&gt;sum &lt;/span&gt;by &lt;span class="o"&gt;(&lt;/span&gt;namespace, pod&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  container_memory_working_set_bytes&lt;span class="o"&gt;{&lt;/span&gt;container!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, image!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In containerd environments the pause container's &lt;code&gt;container&lt;/code&gt; is empty, so &lt;code&gt;container!=""&lt;/code&gt; alone excludes the Pod cgroup, the root cgroup and pause. &lt;code&gt;image!=""&lt;/code&gt; is added because series other than real containers have no &lt;code&gt;image&lt;/code&gt;, and it is insurance for environments where &lt;code&gt;container!=""&lt;/code&gt; alone leaves some series. With CRI-O the pause container appears as &lt;code&gt;container="POD"&lt;/code&gt;, so also add &lt;code&gt;container!="POD"&lt;/code&gt; there.&lt;/p&gt;

&lt;p&gt;The queries from here on use the form that looks at the Pod's cgroup directly, which works in any environment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compare request and actual usage for each Pod&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is the ratio of actual usage to the request. &lt;strong&gt;A Pod above 1 has too small a request&lt;/strong&gt;, and is more likely to be targeted in both eviction and global OOM.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Ratio of actual usage to memory request (above 1 means the request is too small)&lt;/span&gt;
max by &lt;span class="o"&gt;(&lt;/span&gt;namespace, pod&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  container_memory_working_set_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;container&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, pod!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;)&lt;/span&gt;
/
max by &lt;span class="o"&gt;(&lt;/span&gt;namespace, pod&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  kube_pod_resource_request&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;resource&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"memory"&lt;/span&gt;, &lt;span class="nv"&gt;unit&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"bytes"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To see the difference in bytes, do the following. A positive value is "the amount by which the request is exceeded".&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Amount over the request (bytes)&lt;/span&gt;
max by &lt;span class="o"&gt;(&lt;/span&gt;namespace, pod&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  container_memory_working_set_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;container&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, pod!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;)&lt;/span&gt;
-
max by &lt;span class="o"&gt;(&lt;/span&gt;namespace, pod&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  kube_pod_resource_request&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;resource&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"memory"&lt;/span&gt;, &lt;span class="nv"&gt;unit&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"bytes"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Find Pods with no request&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;As noted above, Pods with no request do not appear in &lt;code&gt;kube_pod_resource_request&lt;/code&gt;, so they do not show up in the ratio query. Use &lt;code&gt;unless&lt;/code&gt; to find "Pods that have a usage series but no request series". &lt;strong&gt;These are the BestEffort Pods, that is, the first candidates to be killed in a global OOM.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Pods with no memory request, and their actual usage&lt;/span&gt;
max by &lt;span class="o"&gt;(&lt;/span&gt;namespace, pod&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  container_memory_working_set_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;container&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, pod!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;)&lt;/span&gt;
unless
max by &lt;span class="o"&gt;(&lt;/span&gt;namespace, pod&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  kube_pod_resource_request&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;resource&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"memory"&lt;/span&gt;, &lt;span class="nv"&gt;unit&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"bytes"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Compare the sum of requests and the sum of actual usage per node&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For each node, compare "usage the scheduler knows about (sum of requests)" with "actual usage". &lt;strong&gt;A node where actual usage is larger is one where the scheduler overestimates free space&lt;/strong&gt;, and the risk of global OOM and eviction is high.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;kube_pod_resource_request&lt;/code&gt; and &lt;code&gt;kube_node_status_allocatable&lt;/code&gt; both have a &lt;code&gt;node&lt;/code&gt; label, so these two can be joined without &lt;code&gt;label_replace&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Per node: sum of memory requests / Allocatable&lt;/span&gt;
&lt;span class="nb"&gt;sum &lt;/span&gt;by &lt;span class="o"&gt;(&lt;/span&gt;node&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  max by &lt;span class="o"&gt;(&lt;/span&gt;node, namespace, pod&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
    kube_pod_resource_request&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;resource&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"memory"&lt;/span&gt;, &lt;span class="nv"&gt;unit&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"bytes"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
  &lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="o"&gt;)&lt;/span&gt;
/
max by &lt;span class="o"&gt;(&lt;/span&gt;node&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  kube_node_status_allocatable&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;resource&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"memory"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The actual-usage side comes from cAdvisor and has no &lt;code&gt;node&lt;/code&gt; label, so align labels as in the PromQL above. Using the working set of &lt;code&gt;/kubepods.slice&lt;/code&gt; gives the usage of all Pods on the node without summing per Pod.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Actual usage of all Pods on the node / Allocatable&lt;/span&gt;
max by &lt;span class="o"&gt;(&lt;/span&gt;node&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  container_memory_working_set_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;~&lt;span class="s2"&gt;"/kubepods|/kubepods.slice"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;*&lt;/span&gt; on&lt;span class="o"&gt;(&lt;/span&gt;instance&lt;span class="o"&gt;)&lt;/span&gt; group_left&lt;span class="o"&gt;(&lt;/span&gt;node&lt;span class="o"&gt;)&lt;/span&gt; kubelet_node_name
&lt;span class="o"&gt;)&lt;/span&gt;
/
max by &lt;span class="o"&gt;(&lt;/span&gt;node&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
  kube_node_status_allocatable&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;resource&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"memory"&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Put these two side by side and look for nodes where the lower query (actual usage) exceeds the upper one (requests).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Derive a recommended request from past usage&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Decide the request not from an instantaneous value but from usage over a period. Memory, unlike CPU, is a resource that cannot be reclaimed (incompressible), so &lt;strong&gt;it is safer to base it on a value near the peak, not the average&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Maximum memory usage per Pod over the past 7 days (a guide for the request)&lt;/span&gt;
max_over_time&lt;span class="o"&gt;(&lt;/span&gt;
  max by &lt;span class="o"&gt;(&lt;/span&gt;namespace, pod&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
    container_memory_working_set_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;container&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, pod!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
  &lt;span class="o"&gt;)[&lt;/span&gt;7d:5m]
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you do not want to be pulled around by transient spikes, use a quantile. Be aware, though, that if you use this value as the request, the Pod will exceed its request at peak times and become an eviction candidate.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# 95th percentile over the past 7 days&lt;/span&gt;
quantile_over_time&lt;span class="o"&gt;(&lt;/span&gt;0.95,
  max by &lt;span class="o"&gt;(&lt;/span&gt;namespace, pod&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="o"&gt;(&lt;/span&gt;
    container_memory_working_set_bytes&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;container&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;, pod!&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;""&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;
  &lt;span class="o"&gt;)[&lt;/span&gt;7d:5m]
&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Warning:&lt;/strong&gt; The maximum obtained by this query &lt;strong&gt;misses spikes shorter than Prometheus' scrape interval&lt;/strong&gt;. For applications that temporarily allocate a lot of memory right after startup, the real peak may be higher than this value. It is the same reason as in My working set is below the limit, but the container was OOMKilled. Leave some margin, or decide after understanding the application's startup behavior.&lt;/p&gt;

&lt;p&gt;Also, the working set includes active page cache. For containers with a lot of logging or file I/O, using this value as is for the request makes it too large. Compare with &lt;code&gt;container_memory_rss&lt;/code&gt; and decide which to base it on.&lt;/p&gt;



&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Do not base request design on kubectl top
&lt;/h4&gt;

&lt;p&gt;&lt;code&gt;kubectl top pod&lt;/code&gt; is convenient, but it is not suited to designing requests. The reason lies in the design of Metrics Server itself.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What Metrics Server gets from the kubelet is the value of the &lt;code&gt;/metrics/resource&lt;/code&gt; endpoint, and for memory it is &lt;strong&gt;the working set&lt;/strong&gt; (&lt;a href="https://github.com/kubernetes-sigs/metrics-server/blob/v0.8.1/FAQ.md" rel="noopener noreferrer"&gt;"Memory is reported as the working set at the instant the metric was collected"&lt;/a&gt;). It is the same value as &lt;code&gt;container_memory_working_set_bytes&lt;/code&gt; seen through Prometheus, so it includes page cache in the same way&lt;/li&gt;
&lt;li&gt;Metrics Server &lt;strong&gt;keeps only the latest value&lt;/strong&gt;. The default collection interval is 60 seconds (&lt;code&gt;--metric-resolution&lt;/code&gt;) and no history is kept, so you cannot look into past peaks&lt;/li&gt;
&lt;li&gt;Metrics Server's own README states in its use cases &lt;em&gt;"Don't use Metrics Server when you need: ... An accurate source of resource usage metrics"&lt;/em&gt; (&lt;a href="https://github.com/kubernetes-sigs/metrics-server/blob/v0.8.1/README.md#use-cases" rel="noopener noreferrer"&gt;README.md&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use &lt;code&gt;kubectl top&lt;/code&gt; to get a feel for "what is using a lot right now", and when you decide a request, specify a period and check with PromQL for comparing requests and actual usage.&lt;/p&gt;

&lt;h3&gt;
  
  
  Adjust requests automatically with VPA
&lt;/h3&gt;

&lt;p&gt;Instead of running the PromQL in the previous section by hand, you can have the &lt;a href="https://github.com/kubernetes/autoscaler/tree/master/vertical-pod-autoscaler" rel="noopener noreferrer"&gt;Vertical Pod Autoscaler (VPA)&lt;/a&gt; compute recommendations. The idea is the same (derive from quantiles of past usage), and VPA runs it continuously, combined with OOM detection.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  VPA configuration examples and defaults
  &lt;p&gt;VPA consists of three components.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Recommender&lt;/td&gt;
&lt;td&gt;Computes recommendations from usage history and writes them to &lt;code&gt;status.recommendation&lt;/code&gt; of the VPA object&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Updater&lt;/td&gt;
&lt;td&gt;Evicts Pods that have drifted from the recommendation and prompts their recreation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Admission Controller&lt;/td&gt;
&lt;td&gt;Rewrites &lt;code&gt;resources&lt;/code&gt; to the recommendation with a mutating webhook at Pod creation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Start by looking at recommendations with &lt;code&gt;updateMode: Off&lt;/code&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;You would not want Pods to be rebuilt abruptly, so &lt;strong&gt;it is safe to first have it only compute recommendations with &lt;code&gt;updateMode: Off&lt;/code&gt;&lt;/strong&gt;. In this mode VPA does not touch Pods at all and only writes recommendations to the VPA object. The API definition also says "This can be used for a \"dry run\"".&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;autoscaling.k8s.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;VerticalPodAutoscaler&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sample-vpa&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;default&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;targetRef&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;apps/v1&lt;/span&gt;
    &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Deployment&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sample&lt;/span&gt;
  &lt;span class="na"&gt;updatePolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;updateMode&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Off"&lt;/span&gt;          &lt;span class="c1"&gt;# only compute recommendations. Pods are not changed&lt;/span&gt;
  &lt;span class="na"&gt;resourcePolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;containerPolicies&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;containerName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;app&lt;/span&gt;
      &lt;span class="na"&gt;controlledResources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;memory"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;   &lt;span class="c1"&gt;# default is ["cpu", "memory"]&lt;/span&gt;
      &lt;span class="na"&gt;controlledValues&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;RequestsOnly&lt;/span&gt;    &lt;span class="c1"&gt;# default is RequestsAndLimits&lt;/span&gt;
      &lt;span class="na"&gt;minAllowed&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;128Mi&lt;/span&gt;
      &lt;span class="na"&gt;maxAllowed&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;4Gi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Check the recommendation as follows.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="s"&gt;$ kubectl describe vpa sample-vpa&lt;/span&gt;
&lt;span class="nn"&gt;...&lt;/span&gt;
&lt;span class="na"&gt;Status&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;Recommendation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;Container Recommendations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;Container Name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;  &lt;span class="s"&gt;app&lt;/span&gt;
      &lt;span class="na"&gt;Lower Bound&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;Memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;  &lt;span class="s"&gt;262144k&lt;/span&gt;
      &lt;span class="na"&gt;Target&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;Memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;  &lt;span class="m"&gt;367001600&lt;/span&gt;
      &lt;span class="na"&gt;Uncapped Target&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;Memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;  &lt;span class="m"&gt;367001600&lt;/span&gt;
      &lt;span class="na"&gt;Upper Bound&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;Memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;  &lt;span class="s"&gt;524288k&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;Target&lt;/code&gt; is the recommended request. &lt;code&gt;Lower Bound&lt;/code&gt; / &lt;code&gt;Upper Bound&lt;/code&gt; are guides for "below this is not enough" and "above this is excessive" respectively, and the Updater makes Pods outside this range eviction targets.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Switch to automatic application&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Once you confirm the recommendation is reasonable, change &lt;code&gt;updateMode&lt;/code&gt; to switch to automatic application. The modes are:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;code&gt;updateMode&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;Behavior&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Off&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Only computes recommendations. Pods are not changed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Initial&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Applies recommendations &lt;strong&gt;only at Pod creation&lt;/strong&gt;. Running Pods are not changed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Recreate&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Applies at creation, and also updates running Pods by &lt;strong&gt;evicting and recreating&lt;/strong&gt; them&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;InPlaceOrRecreate&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Tries In-Place Resize first, and falls back to recreation if that is not possible. Requires the cluster's &lt;code&gt;InPlacePodVerticalScaling&lt;/code&gt; feature gate&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;InPlace&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Tries only In-Place Resize and &lt;strong&gt;does not evict&lt;/strong&gt;. On failure it leaves it to the kubelet's retry. Requires VPA's own &lt;code&gt;InPlace&lt;/code&gt; feature gate in addition to the above&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Auto&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Deprecated&lt;/strong&gt;. Currently equivalent to &lt;code&gt;Recreate&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;Auto&lt;/code&gt; has been deprecated as of VPA 1.7, and the API comment also says "Use explicit update modes like \"Recreate\", \"Initial\", or \"InPlaceOrRecreate\" instead". &lt;strong&gt;Do not use &lt;code&gt;Auto&lt;/code&gt; in newly written manifests.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This is a configuration example that automatically adjusts only the memory request.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;autoscaling.k8s.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;VerticalPodAutoscaler&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sample-vpa&lt;/span&gt;
  &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;default&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;targetRef&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;apps/v1&lt;/span&gt;
    &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Deployment&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sample&lt;/span&gt;
  &lt;span class="na"&gt;updatePolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;updateMode&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Recreate"&lt;/span&gt;
    &lt;span class="na"&gt;minReplicas&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;2&lt;/span&gt;             &lt;span class="c1"&gt;# do not evict when replicas is below this (the global default is also 2)&lt;/span&gt;
  &lt;span class="na"&gt;resourcePolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;containerPolicies&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;containerName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;app&lt;/span&gt;
      &lt;span class="na"&gt;controlledResources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;memory"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;controlledValues&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;RequestsOnly&lt;/span&gt;
      &lt;span class="na"&gt;minAllowed&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;128Mi&lt;/span&gt;
      &lt;span class="na"&gt;maxAllowed&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;4Gi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With &lt;code&gt;controlledValues: RequestsOnly&lt;/code&gt;, &lt;strong&gt;you can manage the limit yourself and leave only the request to VPA&lt;/strong&gt;. In the context of this article (reducing global OOM and eviction) this is the easy setting to work with. With the default &lt;code&gt;RequestsAndLimits&lt;/code&gt;, both are rewritten while keeping the ratio of request to limit, so note that the limit moves too.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Behavior when an OOM occurs&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For memory, VPA has a mechanism that raises the recommendation when it detects an OOM kill. The defaults are:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setting&lt;/th&gt;
&lt;th&gt;Default&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;oomBumpUpRatio&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;1.2&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;When an OOM kill is detected, record usage at that time multiplied by 1.2 as a sample&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;oomMinBumpUp&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;100Mi&lt;/code&gt; (104857600)&lt;/td&gt;
&lt;td&gt;If the bump is smaller than this, bump up by at least this much&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;memoryAggregationIntervalSeconds&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;86400&lt;/code&gt; (24 hours)&lt;/td&gt;
&lt;td&gt;Record one peak sample per this interval&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;memoryAggregationIntervalCount&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;8&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;How many such intervals to keep. By default the memory recommendation is computed from 24 hours × 8 = &lt;strong&gt;8 days&lt;/strong&gt; of history&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;evictAfterOOMSeconds&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;(unset)&lt;/td&gt;
&lt;td&gt;Make Pods that OOM within this many seconds of starting eviction targets&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;From VPA 1.7 onward, these can be specified per container in &lt;code&gt;containerPolicies&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;    &lt;span class="na"&gt;containerPolicies&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;containerName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;app&lt;/span&gt;
      &lt;span class="na"&gt;controlledResources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;memory"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;oomBumpUpRatio&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1.5"&lt;/span&gt;                    &lt;span class="c1"&gt;# bump up more strongly on OOM&lt;/span&gt;
      &lt;span class="na"&gt;oomMinBumpUp&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;256Mi&lt;/span&gt;
      &lt;span class="na"&gt;memoryAggregationIntervalSeconds&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;3600&lt;/span&gt;   &lt;span class="c1"&gt;# record a peak every hour&lt;/span&gt;
      &lt;span class="na"&gt;memoryAggregationIntervalCount&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;24&lt;/span&gt;       &lt;span class="c1"&gt;# compute from 24 intervals = 1 day of history&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The memory recommendation itself is computed from quantiles of a usage histogram. The Recommender defaults are:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Flag&lt;/th&gt;
&lt;th&gt;Default&lt;/th&gt;
&lt;th&gt;Use&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--target-memory-percentile&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.9&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Quantile used to compute &lt;code&gt;Target&lt;/code&gt; (the recommended request)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--recommendation-lower-bound-memory-percentile&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Quantile used to compute &lt;code&gt;Lower Bound&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--recommendation-upper-bound-memory-percentile&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.95&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Quantile used to compute &lt;code&gt;Upper Bound&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--recommendation-margin-fraction&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.15&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Safety margin added to the computed value (15 %)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So by default, &lt;strong&gt;the 90th percentile of 8 days of peak values plus a 15 % margin&lt;/strong&gt; becomes the recommended request. The idea is the same as the &lt;code&gt;quantile_over_time(0.95, ...)&lt;/code&gt; PromQL introduced in the previous section, and it is easiest to think of VPA as running that continuously, combined with OOM detection.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Warning: The usage VPA looks at is also the working set.&lt;/strong&gt; The Recommender gets usage from the Metrics Server API (&lt;code&gt;metrics.k8s.io/v1beta1&lt;/code&gt;), so as described earlier it derives recommendations from values that include page cache.&lt;/p&gt;

&lt;p&gt;So &lt;strong&gt;for containers with a lot of logging or file I/O, VPA's recommendation comes out larger than the real need&lt;/strong&gt;. For containers where the gap between &lt;code&gt;container_memory_rss&lt;/code&gt; and &lt;code&gt;container_memory_working_set_bytes&lt;/code&gt; is large, it may waste less in the end to cap with &lt;code&gt;maxAllowed&lt;/code&gt; or to decide by hand instead of leaving it to VPA.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;metrics.k8s.io&lt;/code&gt; went GA (&lt;code&gt;v1&lt;/code&gt;) in v1.37 through &lt;a href="https://github.com/kubernetes/enhancements/issues/5207" rel="noopener noreferrer"&gt;KEP-5207&lt;/a&gt;, but on the Metrics Server side it is still unsupported even in v0.9.0, the latest release at the time of writing (September 2026), and is being addressed in &lt;a href="https://github.com/kubernetes-sigs/metrics-server/issues/1786" rel="noopener noreferrer"&gt;issue #1786&lt;/a&gt; and &lt;a href="https://github.com/kubernetes-sigs/metrics-server/pull/1855" rel="noopener noreferrer"&gt;PR #1855&lt;/a&gt;. That is why this article writes &lt;code&gt;v1beta1&lt;/code&gt;. That the referenced value is the working set does not change with the API version.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Operational points when using VPA:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Combining with HPA&lt;/strong&gt;: &lt;strong&gt;do not use HPA and VPA on the same resource at the same time&lt;/strong&gt;. The official docs also state "should not be used with the HPA on the same resource metric (CPU or memory)" (&lt;a href="https://github.com/kubernetes/autoscaler/blob/vertical-pod-autoscaler-1.7.1/vertical-pod-autoscaler/docs/known-limitations.md" rel="noopener noreferrer"&gt;known-limitations.md&lt;/a&gt;). If you split the targets they can be combined, and &lt;strong&gt;VPA for memory and HPA for CPU&lt;/strong&gt; is the officially recommended example. With &lt;code&gt;controlledResources: ["memory"]&lt;/code&gt; as in this article's examples, CPU can be left to HPA&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;Recreate&lt;/code&gt; rebuilds Pods&lt;/strong&gt;: the Updater evicts Pods using the Eviction API, so set a PodDisruptionBudget. It does not evict when replicas is below &lt;code&gt;minReplicas&lt;/code&gt; (the global &lt;code&gt;--min-replicas&lt;/code&gt;, default &lt;code&gt;2&lt;/code&gt;, if not specified in VPA)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;In-Place modes&lt;/strong&gt;: they can change the request without rebuilding the Pod, but require the cluster's &lt;code&gt;InPlacePodVerticalScaling&lt;/code&gt; feature gate&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metrics Server is required&lt;/strong&gt;: if the Recommender cannot get usage, no recommendation is produced&lt;/li&gt;
&lt;/ul&gt;



&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  My working set is below the limit, but the container was OOMKilled
&lt;/h3&gt;

&lt;p&gt;What memcg OOM compares against the limit is &lt;code&gt;memory.current&lt;/code&gt; (&lt;code&gt;container_memory_usage_bytes&lt;/code&gt;), not the working set. Furthermore, &lt;code&gt;memory.current&lt;/code&gt; reaching the limit does not cause an OOM by itself; the kill happens when reclaim fails to bring usage back under the limit. Also, even if your container is under its limit, it can be killed in a node-wide memory exhaustion (global OOM). First use The two OOM kill paths to tell the path.&lt;/p&gt;

&lt;p&gt;If none of the values seems to have reached the limit, suspect a spike shorter than the scrape interval. A spike in memory usage shorter than Prometheus' scrape interval (the global &lt;code&gt;scrape_interval&lt;/code&gt; defaults to 1 minute; see &lt;a href="https://prometheus.io/docs/prometheus/latest/configuration/configuration/" rel="noopener noreferrer"&gt;Configuration&lt;/a&gt;) does not show up in metrics. Batch processing right after startup, or suddenly accepting a very large request, are examples.&lt;/p&gt;

&lt;p&gt;In that case, check values like these:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Emitted by&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;kube_pod_container_status_last_terminated_reason{reason="OOMKilled"}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;kube-state-metrics&lt;/td&gt;
&lt;td&gt;The last termination reason. The value does not count up when OOMs repeat, so from the second time on, detect it by combining with the increase of &lt;code&gt;kube_pod_container_status_restarts_total&lt;/code&gt; (also kube-state-metrics)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_oom_events_total&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;kubelet&lt;/td&gt;
&lt;td&gt;Number of OOM kills. But &lt;strong&gt;when the container terminates the cgroup and its series disappear&lt;/strong&gt;, so it is limited to checking &lt;strong&gt;cases where only a child process was OOM killed while the container continued&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;node_vmstat_oom_kill&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;node-exporter&lt;/td&gt;
&lt;td&gt;Number of OOM kills on the node. It cannot identify the Pod or container, but is not affected by container recreation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Besides metrics, also check &lt;code&gt;Last State&lt;/code&gt; in &lt;code&gt;kubectl describe pod&lt;/code&gt; (exit code 137 means termination by SIGKILL).&lt;/p&gt;
&lt;h3&gt;
  
  
  Node usage computed from MemAvailable is low, but Pods are evicted
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;MemAvailable&lt;/code&gt; counts part of the page cache as "free", but the kubelet is working-set based and treats all of &lt;code&gt;active_file&lt;/code&gt; as "in use". With workloads that do a lot of file I/O or logging, &lt;code&gt;active_file&lt;/code&gt; can grow to several GiB, so even if the &lt;code&gt;MemAvailable&lt;/code&gt;-based utilization looks low, the kubelet side may be approaching its threshold.&lt;/p&gt;

&lt;p&gt;Conversely, on nodes where page cache has not built up much, node-exporter's side comes out higher. &lt;strong&gt;Which is higher depends on the environment&lt;/strong&gt;, so you cannot judge from just one of them. For the breakdown of factors and the PromQL to get the kubelet-side utilization, see Why node-exporter and kubelet usage disagree.&lt;/p&gt;
&lt;h3&gt;
  
  
  kubectl top pod and the Prometheus working set differ
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;kubectl top pod&lt;/code&gt; shows the sum of the containers' working sets obtained from the kubelet (via Metrics Server). Prometheus' &lt;code&gt;container_memory_working_set_bytes&lt;/code&gt; comes from the same source, but it is as stale as the scrape interval, so instantaneous values differ.&lt;/p&gt;
&lt;h3&gt;
  
  
  Memory working set keeps growing
&lt;/h3&gt;

&lt;p&gt;First check whether &lt;code&gt;container_memory_rss&lt;/code&gt; or &lt;code&gt;container_memory_cache&lt;/code&gt; is the one growing.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;container_memory_rss&lt;/code&gt; is growing: suspect a memory leak in the application&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;container_memory_cache&lt;/code&gt; is growing: page cache is accumulating. But it includes not only the cache of files on disk but also tmpfs and shared memory. To tell which is growing, check &lt;code&gt;shmem&lt;/code&gt; in &lt;code&gt;memory.stat&lt;/code&gt; on the node. tmpfs (such as &lt;code&gt;emptyDir&lt;/code&gt; with &lt;code&gt;medium: Memory&lt;/code&gt;) is not reclaimed on nodes without swap, so accumulation leads to OOM kills. &lt;strong&gt;&lt;code&gt;shmem&lt;/code&gt; sits on the anon LRU, so it does not appear in &lt;code&gt;inactive_file&lt;/code&gt; and is not subtracted from the working set&lt;/strong&gt;. That means the increase pushes the working set up as is&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For containers with no memory limit, page cache accumulates with the whole node's memory as the ceiling. Set a memory limit on containers that handle large files.&lt;/p&gt;
&lt;h3&gt;
  
  
  The RSS my application reports differs from the RSS metric
&lt;/h3&gt;

&lt;p&gt;They measure different things. The value called "RSS" in Go, Java, Node.js and so on is &lt;strong&gt;per-process RSS&lt;/strong&gt;, whereas &lt;code&gt;container_memory_rss&lt;/code&gt; is &lt;strong&gt;anonymous pages charged to the container's cgroup&lt;/strong&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Target&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;Backing value&lt;/th&gt;
&lt;th&gt;What it includes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;process.memoryUsage().rss&lt;/code&gt; in Node.js, the &lt;code&gt;RSS&lt;/code&gt; column of &lt;code&gt;ps&lt;/code&gt; and &lt;code&gt;top&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Language runtime or tools such as &lt;code&gt;ps&lt;/code&gt; (per process)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;VmRSS&lt;/code&gt; in &lt;code&gt;/proc/[pid]/status&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;RssAnon&lt;/code&gt; + &lt;code&gt;RssFile&lt;/code&gt; + &lt;code&gt;RssShmem&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_rss&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet; per cgroup)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;anon&lt;/code&gt; in cgroup v2 &lt;code&gt;memory.stat&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Anonymous pages only&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;There are two main differences:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Whether file mappings are included&lt;/strong&gt;: a process's RSS includes the pages of the executable, shared libraries and &lt;code&gt;mmap&lt;/code&gt;ed files that are in physical memory (&lt;code&gt;RssFile&lt;/code&gt;). &lt;code&gt;container_memory_rss&lt;/code&gt; does not. So &lt;strong&gt;comparing the same single process, the RSS the application reports is larger by &lt;code&gt;RssFile&lt;/code&gt; and &lt;code&gt;RssShmem&lt;/code&gt;&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unit of aggregation&lt;/strong&gt;: a process's RSS is for one process. &lt;code&gt;container_memory_rss&lt;/code&gt; is per container cgroup, so it is the total of all processes in the container. Also, a process's RSS counts shared pages in each process, whereas in a cgroup a page is charged to only one cgroup&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Note that JVM &lt;code&gt;Runtime.totalMemory()&lt;/code&gt; and JMX heap usage are not RSS to begin with; they are the heap size. Metaspace, code cache, thread stacks and GC bookkeeping areas are not included, so you cannot compare these values directly with the memory limit.&lt;/p&gt;

&lt;p&gt;Neither value is used in the OOM kill decision. What memcg OOM compares against the limit is &lt;code&gt;memory.current&lt;/code&gt; (&lt;code&gt;container_memory_usage_bytes&lt;/code&gt;). Global OOM occurs when the memory of the whole node is exhausted and does not use any individual container's usage as a threshold.&lt;/p&gt;
&lt;h3&gt;
  
  
  Which metric should I use for node memory utilization
&lt;/h3&gt;

&lt;p&gt;Use them according to purpose.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;th&gt;Metric to look at&lt;/th&gt;
&lt;th&gt;Emitted by&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Get a rough sense of the node's headroom&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;node_memory_MemAvailable_bytes&lt;/code&gt; / &lt;code&gt;node_memory_MemTotal_bytes&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;node-exporter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Judge whether eviction will occur&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;container_memory_working_set_bytes{id="/"}&lt;/code&gt; / &lt;code&gt;kube_node_status_capacity{resource="memory"}&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;cAdvisor (embedded in kubelet) / kube-state-metrics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Judge whether a Pod can be placed&lt;/td&gt;
&lt;td&gt;The sum of &lt;code&gt;kube_pod_resource_request{resource="memory"}&lt;/code&gt; (the sum of requests, not usage)&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;kube-scheduler&lt;/strong&gt; &lt;code&gt;/metrics/resources&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; What emits &lt;code&gt;kube_pod_resource_request&lt;/code&gt; is &lt;strong&gt;kube-scheduler&lt;/strong&gt;, not kube-state-metrics (&lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/scheduler/metrics/resources/resources.go#L54-L71" rel="noopener noreferrer"&gt;resources.go&lt;/a&gt;). kube-state-metrics also has a similarly named &lt;code&gt;kube_pod_container_resource_requests&lt;/code&gt;, but kube-state-metrics itself says in its help text that it recommends using the more accurate kube-scheduler metrics. Unless Prometheus is configured to scrape kube-scheduler's &lt;code&gt;/metrics/resources&lt;/code&gt;, &lt;code&gt;kube_pod_resource_request&lt;/code&gt; is not collected.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h2&gt;
  
  
  Upcoming changes
&lt;/h2&gt;

&lt;p&gt;These are in-progress Kubernetes changes that relate to the behavior described in this article.&lt;/p&gt;
&lt;h3&gt;
  
  
  KEP-2371: cAdvisor-less, CRI-full Container and Pod Stats
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://github.com/kubernetes/enhancements/tree/master/keps/sig-node/2371-cri-pod-container-stats" rel="noopener noreferrer"&gt;KEP-2371&lt;/a&gt; moves the source of container and node metrics from the kubelet's embedded cAdvisor to the container runtime (CRI). Its goal is to remove the situation where the kubelet collects duplicate statistics from both cAdvisor and CRI.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The &lt;code&gt;/metrics/cadvisor&lt;/code&gt; endpoint and the &lt;code&gt;container_*&lt;/code&gt; metric names are kept, and the plan is to change only where the values come from to CRI&lt;/li&gt;
&lt;li&gt;The feature gate &lt;code&gt;PodAndContainerStatsFromCRI&lt;/code&gt; was Alpha in v1.23 and becomes Beta in v1.37, but &lt;strong&gt;even at Beta it is disabled by default&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the definition of the working set and the eviction decision method described in this article will not change for the time being. If this migration becomes the default in the future, the source of values becomes the runtime side, so small differences may arise from differences in how cgroups are read.&lt;/p&gt;
&lt;h3&gt;
  
  
  KEP-2570: Support Memory QoS with cgroups v2
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://github.com/kubernetes/enhancements/tree/master/keps/sig-node/2570-memory-qos" rel="noopener noreferrer"&gt;KEP-2570&lt;/a&gt; reflects &lt;code&gt;requests.memory&lt;/code&gt; into cgroup v2 memory protection settings. When enabled, the memory for the request is protected from reclaim, and throttling kicks in before the limit is reached.&lt;/p&gt;

&lt;p&gt;The feature gate &lt;code&gt;MemoryQoS&lt;/code&gt; has been Alpha (disabled by default) since v1.22 and &lt;strong&gt;became Beta and enabled by default in v1.37&lt;/strong&gt;. However, &lt;strong&gt;just enabling the feature gate does not make &lt;code&gt;requests.memory&lt;/code&gt; reflected in the cgroup&lt;/strong&gt;. To reflect it, the kubelet needs the following settings, both disabled by default as of v1.37.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;memoryReservationPolicy: TieredReservation&lt;/code&gt;: sets &lt;code&gt;requests.memory&lt;/code&gt; into &lt;code&gt;memory.min&lt;/code&gt; for Guaranteed and &lt;code&gt;memory.low&lt;/code&gt; for the rest. A setting added to KubeletConfiguration in v1.36, with a default of &lt;code&gt;None&lt;/code&gt;. With &lt;code&gt;None&lt;/code&gt;, &lt;code&gt;memory.min&lt;/code&gt; is not set&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;memoryThrottlingFactor&lt;/code&gt;: sets &lt;code&gt;memory.high&lt;/code&gt; to &lt;code&gt;requests.memory + factor × (limits.memory - requests.memory)&lt;/code&gt;. Up to v1.36 the default was &lt;code&gt;0.9&lt;/code&gt;, but the feature gate itself was disabled by default then, so it had no actual effect. In v1.37 the default is &lt;code&gt;nil&lt;/code&gt;, and unless you set it explicitly &lt;code&gt;memory.high&lt;/code&gt; is not set&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Both can be confirmed in the &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/staging/src/k8s.io/kubelet/config/v1beta1/types.go#L899-L917" rel="noopener noreferrer"&gt;KubeletConfiguration type definition&lt;/a&gt;.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;    &lt;span class="c"&gt;// MemoryThrottlingFactor specifies the factor multiplied by the memory limit or node allocatable memory&lt;/span&gt;
    &lt;span class="c"&gt;// ...&lt;/span&gt;
    &lt;span class="c"&gt;// Default: nil&lt;/span&gt;
    &lt;span class="n"&gt;MemoryThrottlingFactor&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="kt"&gt;float64&lt;/span&gt; &lt;span class="s"&gt;`json:"memoryThrottlingFactor,omitempty"`&lt;/span&gt;
    &lt;span class="c"&gt;// MemoryReservationPolicy controls how the kubelet applies cgroup v2 memory protection.&lt;/span&gt;
    &lt;span class="c"&gt;// "None" (default): The kubelet does not set memory.min for containers and pods,&lt;/span&gt;
    &lt;span class="c"&gt;// ...&lt;/span&gt;
    &lt;span class="n"&gt;MemoryReservationPolicy&lt;/span&gt; &lt;span class="n"&gt;MemoryReservationPolicy&lt;/span&gt; &lt;span class="s"&gt;`json:"memoryReservationPolicy,omitempty"`&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;So even on a v1.37 cluster, &lt;code&gt;requests.memory&lt;/code&gt; is not reflected in the cgroup, as in the table above. &lt;strong&gt;"Not reflected" is only about the default configuration&lt;/strong&gt;, and it is reflected if you set the two above.&lt;/p&gt;

&lt;p&gt;What is actually written with the default configuration is as in the &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/kuberuntime/kuberuntime_container_linux.go#L207-L218" rel="noopener noreferrer"&gt;implementation&lt;/a&gt;. When &lt;code&gt;memoryReservationPolicy&lt;/code&gt; is &lt;code&gt;None&lt;/code&gt;, &lt;strong&gt;an explicit &lt;code&gt;0&lt;/code&gt; is written&lt;/strong&gt; instead of &lt;code&gt;requests.memory&lt;/code&gt; (equivalent to no protection).&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;memoryRequest&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="m"&gt;0&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;memoryReservationPolicy&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;kubeletconfiginternal&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;TieredReservationMemoryReservationPolicy&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="o"&gt;...&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="n"&gt;unified&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;cm&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Cgroup2MemoryMin&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"0"&lt;/span&gt;
            &lt;span class="n"&gt;unified&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;cm&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Cgroup2MemoryLow&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"0"&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;So with the v1.37 default configuration, a container's cgroup has &lt;code&gt;memory.min = 0&lt;/code&gt;, &lt;code&gt;memory.low = 0&lt;/code&gt; and &lt;code&gt;memory.high&lt;/code&gt; unset (&lt;code&gt;max&lt;/code&gt;), and only &lt;code&gt;memory.max&lt;/code&gt; (= &lt;code&gt;limits.memory&lt;/code&gt;) is in effect. If you use a managed service, check the provider's release notes to see whether these settings get enabled.&lt;/p&gt;
&lt;h4&gt;
  
  
  What the cgroup v2 values become if you enable it
&lt;/h4&gt;

&lt;p&gt;This section works out what gets written to each cgroup v2 file &lt;strong&gt;when you explicitly set &lt;code&gt;memoryThrottlingFactor: 0.9&lt;/code&gt;, the default up to v1.36, and also enable &lt;code&gt;memoryReservationPolicy: TieredReservation&lt;/code&gt;&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;/p&gt;
  kubelet configuration example and the values written for each QoS class
  &lt;p&gt;The kubelet configuration looks like this. &lt;code&gt;MemoryQoS&lt;/code&gt; is enabled by default in v1.37, so you can omit &lt;code&gt;featureGates&lt;/code&gt;, but I include it to make the intent explicit.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# /var/lib/kubelet/config.yaml (KubeletConfiguration)&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kubelet.config.k8s.io/v1beta1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;KubeletConfiguration&lt;/span&gt;
&lt;span class="na"&gt;featureGates&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;MemoryQoS&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;            &lt;span class="c1"&gt;# true by default in v1.37&lt;/span&gt;
&lt;span class="na"&gt;memoryReservationPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;TieredReservation&lt;/span&gt;   &lt;span class="c1"&gt;# default is None&lt;/span&gt;
&lt;span class="na"&gt;memoryThrottlingFactor&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;0.9&lt;/span&gt;                  &lt;span class="c1"&gt;# default in v1.37 is nil (= memory.high is not set)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;Suppose you create the following Pod in this state. Assume the node's Allocatable memory is 16Gi.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Pod&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sample&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;app&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nginx&lt;/span&gt;
    &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;requests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;256Mi&lt;/span&gt;
      &lt;span class="na"&gt;limits&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;1Gi&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;code&gt;requests.memory&lt;/code&gt; and &lt;code&gt;limits.memory&lt;/code&gt; differ, so this Pod's QoS class is Burstable. The following values are written to the container's cgroup.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;File&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;th&gt;Calculation&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;memory.max&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;1073741824&lt;/code&gt; (1Gi)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;limits.memory&lt;/code&gt; as is (set as before, independent of MemoryQoS)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;memory.low&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;268435456&lt;/code&gt; (256Mi)&lt;/td&gt;
&lt;td&gt;Burstable, so &lt;code&gt;requests.memory&lt;/code&gt; goes to &lt;code&gt;memory.low&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;memory.min&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;requests.memory&lt;/code&gt; goes in only for Guaranteed. Otherwise an explicit &lt;code&gt;0&lt;/code&gt; is written&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;memory.high&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;993210368&lt;/code&gt; (about 947 MiB)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;floor((256Mi + (1Gi - 256Mi) × 0.9) / 4096) × 4096&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code&gt;memory.high&lt;/code&gt; calculation is rounded down to a multiple of the page size (4 KiB), as in the &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/kuberuntime/kuberuntime_container_linux.go#L238-L258" rel="noopener noreferrer"&gt;implementation&lt;/a&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight go"&gt;&lt;code&gt;            &lt;span class="n"&gt;memoryHigh&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kt"&gt;int64&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;math&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Floor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="kt"&gt;float64&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memoryRequest&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;
                    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kt"&gt;float64&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memoryLimitVal&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="kt"&gt;float64&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;memoryRequest&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="kt"&gt;float64&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;memoryThrottlingFactor&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="kt"&gt;float64&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;defaultPageSize&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;defaultPageSize&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So this container is &lt;strong&gt;throttled with strong reclaim pressure at about 947 MiB, before it reaches 1Gi and gets OOM killed&lt;/strong&gt;. Up to 256Mi it is also less likely to be reclaimed thanks to &lt;code&gt;memory.low&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Organized by QoS class, it looks like this. What matters is that even with the same &lt;code&gt;requests&lt;/code&gt; / &lt;code&gt;limits&lt;/code&gt;, the files that get set change with the QoS class.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;QoS class&lt;/th&gt;
&lt;th&gt;&lt;code&gt;memory.min&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;memory.low&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;memory.high&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Guaranteed (&lt;code&gt;requests&lt;/code&gt; = &lt;code&gt;limits&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;requests.memory&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Not set&lt;/strong&gt; (stays &lt;code&gt;max&lt;/code&gt;; excluded when &lt;code&gt;requests&lt;/code&gt; and &lt;code&gt;limits&lt;/code&gt; are equal)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Burstable (&lt;code&gt;requests&lt;/code&gt; &amp;lt; &lt;code&gt;limits&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;requests.memory&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;floor((req + (lim − req) × factor) / page size) × page size&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Burstable (no &lt;code&gt;limits&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;requests.memory&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Same formula using &lt;strong&gt;the node's Allocatable&lt;/strong&gt; instead of &lt;code&gt;limits&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BestEffort (no &lt;code&gt;requests&lt;/code&gt;/&lt;code&gt;limits&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;floor(Allocatable × factor / page size) × page size&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For a BestEffort Pod on the same node as the example above (Allocatable 16Gi), &lt;code&gt;memory.high&lt;/code&gt; is &lt;code&gt;floor(16Gi × 0.9 / 4096) × 4096 = 15461879808&lt;/code&gt; (about 14.4 GiB). &lt;strong&gt;&lt;code&gt;memory.high&lt;/code&gt; is also set on containers with no limit&lt;/strong&gt;, so enabling &lt;code&gt;memoryThrottlingFactor&lt;/code&gt; changes how page cache accumulates as well.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;memory.min&lt;/code&gt; and &lt;code&gt;memory.low&lt;/code&gt; are set not only on the container cgroup but also on higher levels (&lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/cm/qos_container_manager_linux.go#L332-L369" rel="noopener noreferrer"&gt;qos_container_manager_linux.go&lt;/a&gt;). The kernel evaluates protection by walking up through ancestor cgroups, so the same protection is needed above as well.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;&lt;code&gt;memory.min&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;memory.low&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/kubepods.slice&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Sum of Guaranteed requests &lt;strong&gt;+&lt;/strong&gt; sum of Burstable requests&lt;/td&gt;
&lt;td&gt;Sum of Burstable requests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;/kubepods.slice/kubepods-burstable.slice&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Sum of Burstable requests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pod cgroup&lt;/td&gt;
&lt;td&gt;The Pod's request if Guaranteed&lt;/td&gt;
&lt;td&gt;The Pod's request if Burstable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Container cgroup&lt;/td&gt;
&lt;td&gt;As in the table above&lt;/td&gt;
&lt;td&gt;As in the table above&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;You can check how much is protected across the whole node with the following metrics the kubelet exposes (both ALPHA as of v1.37).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Emitted by&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;kubelet_memory_qos_node_memory_min_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;kubelet&lt;/td&gt;
&lt;td&gt;Total reserved as &lt;code&gt;memory.min&lt;/code&gt; for Guaranteed Pods. &lt;strong&gt;The amount of memory the kernel will never reclaim&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;kubelet_memory_qos_node_memory_low_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;kubelet&lt;/td&gt;
&lt;td&gt;Total reserved as &lt;code&gt;memory.low&lt;/code&gt; for Burstable Pods&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Warning:&lt;/strong&gt; If you enable &lt;code&gt;memoryReservationPolicy: TieredReservation&lt;/code&gt;, the sum of Guaranteed Pods' requests is hard-reserved on the node as &lt;code&gt;memory.min&lt;/code&gt;. &lt;strong&gt;The kernel cannot reclaim the &lt;code&gt;memory.min&lt;/code&gt; portion&lt;/strong&gt;, so on nodes with many Guaranteed Pods with oversized requests, the reclaimable memory shrinks and the node may become less stable. The reason the default of &lt;code&gt;memoryReservationPolicy&lt;/code&gt; is &lt;code&gt;None&lt;/code&gt; is also written in the type definition comment: "This is the default to maintain node stability by preventing \"locked\" memory."&lt;/p&gt;

&lt;p&gt;If you are considering enabling it, the precondition is to first right-size requests with Align memory request with actual usage.&lt;/p&gt;



&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Hands-on verification with kind
&lt;/h2&gt;

&lt;p&gt;These two sections reproduce the behavior described above on a local kind cluster.&lt;/p&gt;

&lt;h3&gt;
  
  
  Verifying it with kind
&lt;/h3&gt;

&lt;p&gt;&lt;/p&gt;
  Steps to bring up a local v1.37 cluster with kind and verify
  &lt;p&gt;You can verify everything so far on your machine by bringing up a v1.37 cluster locally with kind. The following configuration file creates an environment with MemoryQoS enabled.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# kind-memqos.yaml&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Cluster&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kind.x-k8s.io/v1alpha4&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;memqos&lt;/span&gt;
&lt;span class="na"&gt;nodes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;control-plane&lt;/span&gt;
  &lt;span class="c1"&gt;# the kind default image changes with the version, so pin it explicitly&lt;/span&gt;
  &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5&lt;/span&gt;
  &lt;span class="na"&gt;kubeadmConfigPatches&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;kind: KubeletConfiguration&lt;/span&gt;
    &lt;span class="s"&gt;featureGates:&lt;/span&gt;
      &lt;span class="s"&gt;MemoryQoS: true&lt;/span&gt;
    &lt;span class="s"&gt;memoryReservationPolicy: TieredReservation&lt;/span&gt;
    &lt;span class="s"&gt;memoryThrottlingFactor: 0.9&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you omit &lt;code&gt;image&lt;/code&gt;, the default image embedded in kind is used, but that changes with each kind version. The digest above is the same as the default of kind v0.33.0; to try a different Kubernetes version, replace it with the digest listed in the &lt;a href="https://github.com/kubernetes-sigs/kind/releases" rel="noopener noreferrer"&gt;kind release notes&lt;/a&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;kind create cluster &lt;span class="nt"&gt;--config&lt;/span&gt; kind-memqos.yaml

&lt;span class="c"&gt;# Check that the settings reached the kubelet&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl get &lt;span class="nt"&gt;--raw&lt;/span&gt; &lt;span class="s2"&gt;"/api/v1/nodes/memqos-control-plane/proxy/configz"&lt;/span&gt; | jq &lt;span class="s1"&gt;'.kubeletconfig | {memoryThrottlingFactor, memoryReservationPolicy, featureGates}'&lt;/span&gt;
&lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="s2"&gt;"memoryThrottlingFactor"&lt;/span&gt;: 0.9,
  &lt;span class="s2"&gt;"memoryReservationPolicy"&lt;/span&gt;: &lt;span class="s2"&gt;"TieredReservation"&lt;/span&gt;,
  &lt;span class="s2"&gt;"featureGates"&lt;/span&gt;: &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="s2"&gt;"MemoryQoS"&lt;/span&gt;: &lt;span class="nb"&gt;true&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You read cgroup files by going into the node container. &lt;strong&gt;kind puts the cgroup root at &lt;code&gt;/kubelet.slice/kubelet-kubepods.slice&lt;/code&gt;&lt;/strong&gt;, so note that the path differs from &lt;code&gt;/kubepods.slice&lt;/code&gt; in a real cluster.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Warning: Two behaviors cannot be reproduced in kind.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One is the root cgroup fallback described earlier. A kind node is a container and has its own cgroup namespace, so the &lt;code&gt;/sys/fs/cgroup&lt;/code&gt; visible from inside the node is actually a delegated &lt;strong&gt;non-root&lt;/strong&gt; cgroup. As a result &lt;code&gt;memory.current&lt;/code&gt; exists, and the &lt;code&gt;anon + file&lt;/code&gt; substitution does not happen. The difference from a real machine is as follows.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Inside a kind node (root of the cgroup namespace = actually a non-root cgroup)&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;docker &lt;span class="nb"&gt;exec &lt;/span&gt;memqos-control-plane &lt;span class="nb"&gt;ls&lt;/span&gt; /sys/fs/cgroup/memory.current
/sys/fs/cgroup/memory.current

&lt;span class="c"&gt;# On the real host (the root cgroup of cgroup v2)&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;colima ssh &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="nb"&gt;ls&lt;/span&gt; /sys/fs/cgroup/memory.current
&lt;span class="nb"&gt;ls&lt;/span&gt;: /sys/fs/cgroup/memory.current: No such file or directory
&lt;span class="nv"&gt;$ &lt;/span&gt;colima ssh &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="nb"&gt;ls&lt;/span&gt; /sys/fs/cgroup/memory.stat
/sys/fs/cgroup/memory.stat
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The other is eviction. kind's kubelet overrides &lt;code&gt;evictionHard&lt;/code&gt; with only &lt;code&gt;imagefs.available&lt;/code&gt; / &lt;code&gt;nodefs.available&lt;/code&gt; / &lt;code&gt;nodefs.inodesFree&lt;/code&gt;, so &lt;strong&gt;there is no &lt;code&gt;memory.available&lt;/code&gt; threshold&lt;/strong&gt;. As a result memory-triggered eviction does not occur, and Capacity and Allocatable are the same value. To try eviction, set &lt;code&gt;evictionHard&lt;/code&gt; explicitly.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Identify the cgroup by the real container ID (a glob can pick up the pause side, as explained below)&lt;/span&gt;
&lt;span class="nv"&gt;$ CID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;kubectl get pod sample-burstable &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;jsonpath&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'{.status.containerStatuses[0].containerID}'&lt;/span&gt; | &lt;span class="nb"&gt;sed&lt;/span&gt; &lt;span class="s1"&gt;'s|containerd://||'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nv"&gt;$ &lt;/span&gt;docker &lt;span class="nb"&gt;exec &lt;/span&gt;memqos-control-plane sh &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s2"&gt;"
    d=&lt;/span&gt;&lt;span class="se"&gt;\$&lt;/span&gt;&lt;span class="s2"&gt;(find /sys/fs/cgroup/kubelet.slice -type d -name 'cri-containerd-&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;CID&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;.scope')
    grep . &lt;/span&gt;&lt;span class="se"&gt;\$&lt;/span&gt;&lt;span class="s2"&gt;d/memory.max &lt;/span&gt;&lt;span class="se"&gt;\$&lt;/span&gt;&lt;span class="s2"&gt;d/memory.min &lt;/span&gt;&lt;span class="se"&gt;\$&lt;/span&gt;&lt;span class="s2"&gt;d/memory.low &lt;/span&gt;&lt;span class="se"&gt;\$&lt;/span&gt;&lt;span class="s2"&gt;d/memory.high"&lt;/span&gt;
/sys/fs/cgroup/.../cri-containerd-50fc89e1....scope/memory.max:1073741824
/sys/fs/cgroup/.../cri-containerd-50fc89e1....scope/memory.min:0
/sys/fs/cgroup/.../cri-containerd-50fc89e1....scope/memory.low:268435456
/sys/fs/cgroup/.../cri-containerd-50fc89e1....scope/memory.high:993210368
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These are measured values on a node with Allocatable of 2005512Ki (= 2053644288 bytes) with &lt;code&gt;memoryThrottlingFactor: 0.9&lt;/code&gt; set.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pod&lt;/th&gt;
&lt;th&gt;QoS&lt;/th&gt;
&lt;th&gt;&lt;code&gt;memory.min&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;memory.low&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;memory.high&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;requests: 256Mi&lt;/code&gt; / &lt;code&gt;limits: 1Gi&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Burstable&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;268435456&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;993210368&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;requests: 128Mi&lt;/code&gt; / no limits&lt;/td&gt;
&lt;td&gt;Burstable&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;134217728&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1861701632&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;memory only &lt;code&gt;requests&lt;/code&gt; = &lt;code&gt;limits&lt;/code&gt; = 512Mi&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Burstable&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;536870912&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;max&lt;/strong&gt; (not set)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cpu and memory &lt;code&gt;requests&lt;/code&gt; = &lt;code&gt;limits&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Guaranteed&lt;/td&gt;
&lt;td&gt;536870912&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;max&lt;/strong&gt; (not set)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;nothing specified&lt;/td&gt;
&lt;td&gt;BestEffort&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1848279040&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each &lt;code&gt;memory.high&lt;/code&gt; follows the formula &lt;code&gt;floor((req + (lim − req) × factor) / page size) × page size&lt;/code&gt; described earlier. For Pods with no &lt;code&gt;limits&lt;/code&gt;, the node's Allocatable is used for &lt;code&gt;lim&lt;/code&gt;, and for BestEffort &lt;code&gt;req&lt;/code&gt; is 0.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Warning: Setting only memory to &lt;code&gt;requests&lt;/code&gt; = &lt;code&gt;limits&lt;/code&gt; does not make a Pod Guaranteed.&lt;/strong&gt; Guaranteed requires &lt;strong&gt;both cpu and memory&lt;/strong&gt; to have &lt;code&gt;requests&lt;/code&gt; = &lt;code&gt;limits&lt;/code&gt; in every container. A Pod with only memory matched is Burstable, so &lt;code&gt;requests.memory&lt;/code&gt; goes to &lt;code&gt;memory.low&lt;/code&gt;, not &lt;code&gt;memory.min&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;memory.high&lt;/code&gt; stays unset (&lt;code&gt;max&lt;/code&gt;) not because of the QoS class but on the condition &lt;code&gt;requests.memory == limits.memory&lt;/code&gt; (the &lt;code&gt;memoryRequest != memoryLimitSpec&lt;/code&gt; check in the &lt;a href="https://github.com/kubernetes/kubernetes/blob/v1.37.0/pkg/kubelet/kuberuntime/kuberuntime_container_linux.go#L226" rel="noopener noreferrer"&gt;implementation&lt;/a&gt;), so even a Burstable Pod with only memory matched stays at &lt;code&gt;max&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;                                    memory.min   memory.low   memory.high
Burstable with only memory matched           0    536870912           max
Guaranteed with cpu and memory matched  536870912          0           max
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Also, besides the real container, a Pod has a pause (sandbox) container's &lt;code&gt;cri-containerd-*.scope&lt;/code&gt; alongside it, so a glob for &lt;code&gt;cri-containerd-*.scope&lt;/code&gt; &lt;strong&gt;matches two entries&lt;/strong&gt;. Which comes first depends on the container ID order, so using the first glob result may read the pause side, and &lt;code&gt;memory.low&lt;/code&gt; and &lt;code&gt;memory.high&lt;/code&gt; can look like &lt;code&gt;0&lt;/code&gt; / &lt;code&gt;max&lt;/code&gt;. Identify it with &lt;code&gt;.status.containerStatuses[].containerID&lt;/code&gt; as in the command above.&lt;/p&gt;

&lt;p&gt;Protection is also set on higher levels, not only on the container cgroup. On &lt;code&gt;kubelet-kubepods.slice&lt;/code&gt; (equivalent to &lt;code&gt;/kubepods.slice&lt;/code&gt; in a real cluster), &lt;code&gt;memory.min&lt;/code&gt; is the sum of Guaranteed and Burstable requests and &lt;code&gt;memory.low&lt;/code&gt; is the sum of Burstable requests.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubelet-kubepods.slice            memory.min 1243611136   memory.low 706740224
kubelet-kubepods-burstable.slice  memory.min 0            memory.low 706740224
kubelet-kubepods-besteffort.slice memory.min 0            memory.low 0

1243611136 - 706740224 &lt;span class="o"&gt;=&lt;/span&gt; 536870912 &lt;span class="o"&gt;=&lt;/span&gt; 512Mi  ← matches the Guaranteed Pod&lt;span class="s1"&gt;'s request
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The kubelet metrics described earlier can also be obtained from this endpoint. &lt;code&gt;memory_min_bytes&lt;/code&gt; covers only the Guaranteed part, so note that its value differs from the root cgroup's &lt;code&gt;memory.min&lt;/code&gt; (Guaranteed + Burstable).&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl get &lt;span class="nt"&gt;--raw&lt;/span&gt; &lt;span class="s2"&gt;"/api/v1/nodes/memqos-control-plane/proxy/metrics"&lt;/span&gt; | &lt;span class="nb"&gt;grep &lt;/span&gt;kubelet_memory_qos
kubelet_memory_qos_node_memory_low_bytes 7.06740224e+08
kubelet_memory_qos_node_memory_min_bytes 5.36870912e+08
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Verifying "reclaim comes first even at the limit" with kind
&lt;/h3&gt;

&lt;p&gt;&lt;/p&gt;
  Confirming that no OOM occurs when page cache pins usage at the limit
  &lt;p&gt;The claim from earlier that "&lt;code&gt;memory.current&lt;/code&gt; reaching &lt;code&gt;memory.max&lt;/code&gt; does not cause an OOM" is easy to reproduce with kind. From a container with &lt;code&gt;limits.memory: 256Mi&lt;/code&gt;, write 6 times the limit of data to an on-disk &lt;code&gt;emptyDir&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl &lt;span class="nb"&gt;exec &lt;/span&gt;memtest &lt;span class="nt"&gt;--&lt;/span&gt; sh &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'dd if=/dev/zero of=/disk/big bs=1M count=1536'&lt;/span&gt;
1610612736 bytes &lt;span class="o"&gt;(&lt;/span&gt;1.5GB&lt;span class="o"&gt;)&lt;/span&gt; copied, 1.501464 seconds, 1023.0MB/s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Looking at the container's cgroup, usage is pinned just short of the limit, but &lt;strong&gt;no OOM occurred even once&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;memory.max      268435456   &lt;span class="o"&gt;(&lt;/span&gt;256Mi&lt;span class="o"&gt;)&lt;/span&gt;
memory.high     248299520   ← floor&lt;span class="o"&gt;((&lt;/span&gt;64Mi + &lt;span class="o"&gt;(&lt;/span&gt;256Mi - 64Mi&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; 0.9&lt;span class="o"&gt;)&lt;/span&gt; / 4096&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; 4096
memory.current  246484992
memory.peak     249049088

&lt;span class="nt"&gt;---&lt;/span&gt; memory.events &lt;span class="nt"&gt;---&lt;/span&gt;
low 0
high 2666        ← number of &lt;span class="nb"&gt;times &lt;/span&gt;it was throttled &lt;span class="k"&gt;for &lt;/span&gt;exceeding memory.high
max 0            ← memory.max was not reached
oom 0
oom_kill 0       ← no OOM &lt;span class="nb"&gt;kill &lt;/span&gt;occurred
oom_group_kill 0

&lt;span class="nt"&gt;---&lt;/span&gt; memory.stat &lt;span class="o"&gt;(&lt;/span&gt;excerpt&lt;span class="o"&gt;)&lt;/span&gt; &lt;span class="nt"&gt;---&lt;/span&gt;
file           235982848
inactive_file  235937792   ← nearly all of it is reclaimable inactive page cache
active_file        45056
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At this point &lt;strong&gt;the working set is &lt;code&gt;246484992 - 235937792 = 10547200&lt;/code&gt; (about 10 MiB)&lt;/strong&gt;. The relationship is observed directly: even when usage is pinned at the limit, if the contents are reclaimable page cache, the working set stays small and no OOM occurs.&lt;/p&gt;

&lt;p&gt;On the other hand, doing the same with anonymous pages causes an immediate OOM kill. &lt;code&gt;tail /dev/zero&lt;/code&gt; keeps allocating memory until it finishes reading its input, and on a node with no swap it cannot be reclaimed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl run oomtest &lt;span class="nt"&gt;--image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;busybox:1.37 &lt;span class="nt"&gt;--restart&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;Never &lt;span class="se"&gt;\&lt;/span&gt;
    &lt;span class="nt"&gt;--overrides&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'{"spec":{"containers":[{"name":"app","image":"busybox:1.37",
      "command":["sh","-c","tail /dev/zero"],
      "resources":{"limits":{"memory":"64Mi"}}}]}}'&lt;/span&gt;

&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl get pod oomtest &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;jsonpath&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'{.status.containerStatuses[0].state.terminated}'&lt;/span&gt;
&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"exitCode"&lt;/span&gt;:137,&lt;span class="s2"&gt;"reason"&lt;/span&gt;:&lt;span class="s2"&gt;"OOMKilled"&lt;/span&gt;,...&lt;span class="o"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The kernel log has a line containing &lt;code&gt;oom_memcg=&lt;/code&gt;. As described earlier, cAdvisor's parser matches this format, so the cgroup can be extracted, and as a result no &lt;code&gt;SystemOOM&lt;/code&gt; event is recorded.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;oom-kill:constraint&lt;span class="o"&gt;=&lt;/span&gt;CONSTRAINT_MEMCG,nodemask&lt;span class="o"&gt;=(&lt;/span&gt;null&lt;span class="o"&gt;)&lt;/span&gt;,cpuset&lt;span class="o"&gt;=&lt;/span&gt;cri-containerd-9be8306a....scope,
&lt;span class="nv"&gt;mems_allowed&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0,oom_memcg&lt;span class="o"&gt;=&lt;/span&gt;/docker/.../kubelet-kubepods-burstable-pod....slice,
&lt;span class="nv"&gt;task_memcg&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/docker/.../cri-containerd-9be8306a....scope,task&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;tail&lt;/span&gt;,pid&lt;span class="o"&gt;=&lt;/span&gt;4867,uid&lt;span class="o"&gt;=&lt;/span&gt;0

Memory cgroup out of memory: Killed process 4867 &lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;tail&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt; total-vm:68880kB,
anon-rss:64768kB, file-rss:1280kB, shmem-rss:0kB, UID:0 pgtables:176kB oom_score_adj:984
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;kubectl get events &lt;span class="nt"&gt;-A&lt;/span&gt; &lt;span class="nt"&gt;--field-selector&lt;/span&gt; &lt;span class="nv"&gt;reason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;SystemOOM
No resources found
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note that after an OOM kill &lt;strong&gt;the container's cgroup disappears, so the &lt;code&gt;container_oom_events_total&lt;/code&gt; series itself goes away&lt;/strong&gt;. It is the same even if you do not let it be recreated with &lt;code&gt;restartPolicy: Never&lt;/code&gt;. To trace the number of OOMs afterward, use &lt;code&gt;node_vmstat_oom_kill&lt;/code&gt; (per node) or &lt;code&gt;kube_pod_container_status_last_terminated_reason&lt;/code&gt;.&lt;/p&gt;



&lt;br&gt;
&lt;p&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing thoughts
&lt;/h2&gt;

&lt;p&gt;The takeaway of this article is that the value you casually call "memory usage" is actually several different ones, namely the working set, RSS, &lt;code&gt;memory.current&lt;/code&gt; and &lt;code&gt;MemAvailable&lt;/code&gt;, and that the value that triggers an OOM kill is different from the value that triggers an eviction.&lt;/p&gt;

&lt;p&gt;I often get questions like "the container's memory usage keeps growing, so isn't it leaking memory?" and "it is &lt;code&gt;OOMKilled&lt;/code&gt; but I don't know why", and I kept giving similar explanations each time. Eviction and QoS classes are in the official documentation, and the cgroup v2 specification is in the kernel documentation, but &lt;strong&gt;I could not find anything that explains in one place what is actually used as a container's memory usage, and why OOM kills and evictions happen as a result&lt;/strong&gt;, so I decided to write it myself.&lt;/p&gt;

&lt;p&gt;Three points are useful to keep in mind when investigating:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What you see with &lt;code&gt;kubectl top&lt;/code&gt; or &lt;code&gt;container_memory_working_set_bytes&lt;/code&gt; is the working set, and it includes active page cache&lt;/li&gt;
&lt;li&gt;What memcg OOM compares against the limit is &lt;code&gt;memory.current&lt;/code&gt; (= &lt;code&gt;container_memory_usage_bytes&lt;/code&gt;), not the working set. And it does not occur just by reaching the limit; the kill happens when reclaim fails to bring usage back under it&lt;/li&gt;
&lt;li&gt;What triggers eviction is free memory computed from the node's working set, not &lt;code&gt;MemAvailable&lt;/code&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Components and versions
&lt;/h3&gt;

&lt;p&gt;This article assumes cgroup v2. With cgroup v1, the mapping between metrics and &lt;code&gt;memory.stat&lt;/code&gt; fields is different.&lt;/p&gt;

&lt;p&gt;The statements in this article target the following versions.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Version&lt;/th&gt;
&lt;th&gt;Repository&lt;/th&gt;
&lt;th&gt;Role in this article&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Linux kernel&lt;/td&gt;
&lt;td&gt;v6.12&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/torvalds/linux" rel="noopener noreferrer"&gt;https://github.com/torvalds/linux&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;cgroup v2 memory control and the OOM killer itself; the OOM message format in kernel logs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kubernetes (kubelet)&lt;/td&gt;
&lt;td&gt;v1.37.0&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/kubernetes/kubernetes" rel="noopener noreferrer"&gt;https://github.com/kubernetes/kubernetes&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Eviction decisions, &lt;code&gt;oom_score_adj&lt;/code&gt;, the &lt;code&gt;SystemOOM&lt;/code&gt; event, counting for &lt;code&gt;container_oom_events_total&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kubernetes (kube-scheduler)&lt;/td&gt;
&lt;td&gt;v1.37.0&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/kubernetes/kubernetes" rel="noopener noreferrer"&gt;https://github.com/kubernetes/kubernetes&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Exposing &lt;code&gt;kube_pod_resource_request&lt;/code&gt; / &lt;code&gt;kube_pod_resource_limit&lt;/code&gt; (&lt;code&gt;/metrics/resources&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;cAdvisor&lt;/td&gt;
&lt;td&gt;v0.57.0&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/google/cadvisor" rel="noopener noreferrer"&gt;https://github.com/google/cadvisor&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Embedded in the kubelet; exposes &lt;code&gt;container_*&lt;/code&gt; metrics; computes the working set&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;containerd&lt;/td&gt;
&lt;td&gt;v2.3.5&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/containerd/containerd" rel="noopener noreferrer"&gt;https://github.com/containerd/containerd&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Decides that a container's termination reason is &lt;code&gt;OOMKilled&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;node-exporter&lt;/td&gt;
&lt;td&gt;v1.12.1&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/prometheus/node_exporter" rel="noopener noreferrer"&gt;https://github.com/prometheus/node_exporter&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Exposes &lt;code&gt;node_memory_*&lt;/code&gt; (&lt;code&gt;/proc/meminfo&lt;/code&gt;) and &lt;code&gt;node_vmstat_oom_kill&lt;/code&gt; (&lt;code&gt;/proc/vmstat&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;kube-state-metrics&lt;/td&gt;
&lt;td&gt;v2.20.0&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/kubernetes/kube-state-metrics" rel="noopener noreferrer"&gt;https://github.com/kubernetes/kube-state-metrics&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Exposes &lt;code&gt;kube_node_status_capacity&lt;/code&gt; and &lt;code&gt;kube_pod_container_status_*&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Metrics Server&lt;/td&gt;
&lt;td&gt;v0.8.1&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/kubernetes-sigs/metrics-server" rel="noopener noreferrer"&gt;https://github.com/kubernetes-sigs/metrics-server&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Source of the values &lt;code&gt;kubectl top&lt;/code&gt; reads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vertical Pod Autoscaler&lt;/td&gt;
&lt;td&gt;1.7.1&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/kubernetes/autoscaler" rel="noopener noreferrer"&gt;https://github.com/kubernetes/autoscaler&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Computing and automatically applying request/limit recommendations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;kind&lt;/td&gt;
&lt;td&gt;v0.33.0&lt;/td&gt;
&lt;td&gt;&lt;a href="https://github.com/kubernetes-sigs/kind" rel="noopener noreferrer"&gt;https://github.com/kubernetes-sigs/kind&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Used to verify v1.37 behavior locally&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; I assume containerd as the container runtime. With other runtimes such as CRI-O the meaning of cgroup v2 values is the same, but the &lt;code&gt;OOMKilled&lt;/code&gt; detection logic and cgroup path naming differ.&lt;/p&gt;

&lt;p&gt;I also assume the cluster has Prometheus, node-exporter and kube-state-metrics installed and that you can run PromQL. These are not part of Kubernetes itself, so if you do not have them, substitute &lt;code&gt;kubectl top&lt;/code&gt; or read &lt;code&gt;memory.stat&lt;/code&gt; directly on the node.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h3&gt;
  
  
  Links
&lt;/h3&gt;

&lt;p&gt;Listed in the order they appear in the article.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://docs.kernel.org/admin-guide/cgroup-v2.html#memory" rel="noopener noreferrer"&gt;Memory Resource Controller (cgroup v2)&lt;/a&gt; (definitions of &lt;code&gt;memory.current&lt;/code&gt; / &lt;code&gt;memory.max&lt;/code&gt; / &lt;code&gt;memory.stat&lt;/code&gt; and other files)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://kubernetes.io/docs/concepts/architecture/cgroups/" rel="noopener noreferrer"&gt;About cgroup v2&lt;/a&gt; (cgroup v2 from Kubernetes' point of view)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/" rel="noopener noreferrer"&gt;Resource Management for Pods and Containers&lt;/a&gt; (basics of requests and limits)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://kubernetes.io/docs/tasks/administer-cluster/reserve-compute-resources/" rel="noopener noreferrer"&gt;Reserve Compute Resources for System Daemons&lt;/a&gt; (calculation of Allocatable)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/node-pressure-eviction/" rel="noopener noreferrer"&gt;Node-pressure Eviction&lt;/a&gt; (overview of eviction: signals, thresholds, selection of target Pods)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.kernel.org/admin-guide/mm/concepts.html#oom-killer" rel="noopener noreferrer"&gt;Concepts overview - OOM killer&lt;/a&gt; (conditions and role of the OOM killer invocation)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/node-pressure-eviction/#node-out-of-memory-behavior" rel="noopener noreferrer"&gt;Node out of memory behavior&lt;/a&gt; (behavior when eviction is not in time)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://kubernetes.io/docs/concepts/workloads/pods/pod-qos/" rel="noopener noreferrer"&gt;Pod Quality of Service Classes&lt;/a&gt; (how QoS classes are determined)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.kernel.org/filesystems/proc.html" rel="noopener noreferrer"&gt;proc.rst - oom_score / oom_score_adj&lt;/a&gt; (the score used to select the process to kill)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://docs.kernel.org/admin-guide/sysctl/vm.html" rel="noopener noreferrer"&gt;sysctl/vm.rst&lt;/a&gt; (OOM-related sysctls such as &lt;code&gt;panic_on_oom&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://kubernetes.io/docs/reference/instrumentation/understand-psi-metrics/" rel="noopener noreferrer"&gt;Understand Pressure Stall Information (PSI) Metrics&lt;/a&gt; (detecting waits for memory allocation)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/kubernetes-sigs/metrics-server/blob/v0.8.1/FAQ.md" rel="noopener noreferrer"&gt;Metrics Server FAQ&lt;/a&gt; (definition and limits of the values &lt;code&gt;kubectl top&lt;/code&gt; shows)&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/kubernetes/autoscaler/tree/master/vertical-pod-autoscaler" rel="noopener noreferrer"&gt;Vertical Pod Autoscaler&lt;/a&gt; (recommending requests and limits from usage)&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Metrics covered in this article
&lt;/h3&gt;

&lt;p&gt;Grouped by emitting component. The takeaway of this article is that &lt;strong&gt;even metrics representing the same "memory usage" have different definitions if they come from different emitters&lt;/strong&gt;, so when investigating, first check which component emits the value.&lt;/p&gt;

&lt;h4&gt;
  
  
  cAdvisor (kubelet &lt;code&gt;/metrics/cadvisor&lt;/code&gt;)
&lt;/h4&gt;

&lt;p&gt;cAdvisor embedded in the kubelet reads cgroups and exposes them. You do not need to run a separate cAdvisor.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Backing value in cgroup v2&lt;/th&gt;
&lt;th&gt;Summary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_usage_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;memory.current&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Usage including everything. &lt;strong&gt;What memcg OOM is evaluated against&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_working_set_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;memory.current - inactive_file&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The value Kubernetes treats as usage. The basis for &lt;code&gt;kubectl top&lt;/code&gt; and eviction decisions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_rss&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;anon&lt;/code&gt; in &lt;code&gt;memory.stat&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Amount of anonymous pages. Its scope differs from true RSS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_cache&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;file&lt;/code&gt; in &lt;code&gt;memory.stat&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Page cache size (includes &lt;code&gt;shmem&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_total_active_file_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;active_file&lt;/code&gt; in &lt;code&gt;memory.stat&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Active page cache. Included in the working set&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_total_inactive_file_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;inactive_file&lt;/code&gt; in &lt;code&gt;memory.stat&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Inactive page cache. Subtracted from the working set&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_mapped_file&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;file_mapped&lt;/code&gt; in &lt;code&gt;memory.stat&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Amount of &lt;code&gt;mmap&lt;/code&gt;ed files&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_spec_memory_limit_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;memory.max&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The container's memory limit. Exposed as &lt;strong&gt;0&lt;/strong&gt; when &lt;code&gt;memory.max&lt;/code&gt; is &lt;code&gt;max&lt;/code&gt; (no limit)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_pressure_memory_waiting_seconds_total&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;total&lt;/code&gt; of &lt;code&gt;some&lt;/code&gt; in &lt;code&gt;memory.pressure&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Cumulative time (seconds) some processes were waiting for memory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_pressure_memory_stalled_seconds_total&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;total&lt;/code&gt; of &lt;code&gt;full&lt;/code&gt; in &lt;code&gt;memory.pressure&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Cumulative time (seconds) all processes were stalled waiting for memory&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The whole node can be selected with &lt;code&gt;id="/"&lt;/code&gt;, all Pods on the node with &lt;code&gt;id=~"/kubepods|/kubepods.slice"&lt;/code&gt;, and a single Pod with &lt;code&gt;container="", image="", pod!=""&lt;/code&gt;. Only &lt;code&gt;id="/"&lt;/code&gt; has no &lt;code&gt;memory.current&lt;/code&gt;, so its usage is substituted with &lt;code&gt;anon + file&lt;/code&gt; from &lt;code&gt;memory.stat&lt;/code&gt; (kernel memory is not included). For details see Values the kubelet uses for eviction decisions.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Emitted by&lt;/th&gt;
&lt;th&gt;Summary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_oom_events_total&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;kubelet&lt;/strong&gt; (the name says &lt;code&gt;container_&lt;/code&gt; but it is not cAdvisor)&lt;/td&gt;
&lt;td&gt;Number of OOM kills parsed from &lt;code&gt;oom-kill:&lt;/code&gt; lines in the kernel log. Exposed on the same &lt;code&gt;/metrics/cadvisor&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h4&gt;
  
  
  kubelet (&lt;code&gt;/metrics/resource&lt;/code&gt;)
&lt;/h4&gt;

&lt;p&gt;An endpoint that exposes only CPU, memory and swap usage per container, per Pod and per node. It is a different path from &lt;code&gt;/metrics/cadvisor&lt;/code&gt;, so to scrape it with Prometheus you need to add configuration.&lt;/p&gt;

&lt;p&gt;Metrics Server gets values from here, and the official documentation also says "The metrics-server fetches resource metrics from the kubelets" (&lt;a href="https://kubernetes.io/docs/tasks/debug/debug-cluster/resource-metrics-pipeline/" rel="noopener noreferrer"&gt;Resource metrics pipeline&lt;/a&gt;). It is not exclusive to Metrics Server, though, and Metrics Server's own README &lt;strong&gt;directs you to scrape this endpoint directly for monitoring purposes&lt;/strong&gt; ("In such cases please collect metrics from Kubelet &lt;code&gt;/metrics/resource&lt;/code&gt; endpoint directly.").&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Summary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;container_memory_working_set_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Per-container working set (the same value as the metric of the same name on &lt;code&gt;/metrics/cadvisor&lt;/code&gt;)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;pod_memory_working_set_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Per-Pod&lt;/strong&gt; working set. Obtained from the Pod's cgroup&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;node_memory_working_set_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;Per-node&lt;/strong&gt; working set&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h4&gt;
  
  
  kubelet (&lt;code&gt;/metrics&lt;/code&gt;)
&lt;/h4&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Summary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;kubelet_node_name&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;A gauge whose value is always &lt;code&gt;1&lt;/code&gt; (the info-metric pattern). The node name is in the &lt;code&gt;node&lt;/code&gt; label, so you can bring &lt;code&gt;node&lt;/code&gt; into the &lt;code&gt;container_*&lt;/code&gt; metrics, which have no node label, with &lt;code&gt;on(instance) group_left(node)&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;kubelet_memory_qos_node_memory_min_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Total reserved as &lt;code&gt;memory.min&lt;/code&gt; for Guaranteed Pods (when MemoryQoS is enabled. ALPHA)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;kubelet_memory_qos_node_memory_low_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Total reserved as &lt;code&gt;memory.low&lt;/code&gt; for Burstable Pods (when MemoryQoS is enabled. ALPHA)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h4&gt;
  
  
  node-exporter
&lt;/h4&gt;

&lt;p&gt;Turns the contents of &lt;code&gt;/proc&lt;/code&gt; directly into metrics. It knows nothing of Kubernetes concepts.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;Summary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;node_memory_MemTotal_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/proc/meminfo&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Total physical memory on the node&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;node_memory_MemAvailable_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/proc/meminfo&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The kernel's estimate of allocatable memory. &lt;strong&gt;Counts reclaimable page cache as free&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;node_memory_Cached_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/proc/meminfo&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Page cache size (includes tmpfs / shmem)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;node_memory_Shmem_bytes&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/proc/meminfo&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The tmpfs / shared memory part of the above&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;node_vmstat_oom_kill&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/proc/vmstat&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Number of OOM kills on the node. Includes both memcg OOM and global OOM&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h4&gt;
  
  
  kube-state-metrics
&lt;/h4&gt;

&lt;p&gt;Converts Kubernetes API objects (spec / status) into metrics. &lt;strong&gt;It does not handle actual resource usage.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Summary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;kube_node_status_capacity{resource="memory"}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The node's Capacity. The denominator for evaluating eviction thresholds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;kube_node_status_allocatable{resource="memory"}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The amount assignable to Pods. The upper bound the scheduler uses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;kube_pod_container_status_last_terminated_reason&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The last termination reason (&lt;code&gt;reason="OOMKilled"&lt;/code&gt; detects an OOM kill). &lt;strong&gt;No series appears for a container that has never terminated&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;kube_pod_container_status_restarts_total&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Container restart count&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;kube_pod_container_resource_requests&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The request written in the manifest (per container)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;kube_pod_container_resource_limits&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The limit written in the manifest (per container)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h4&gt;
  
  
  kube-scheduler (&lt;code&gt;/metrics/resources&lt;/code&gt;)
&lt;/h4&gt;

&lt;p&gt;Exposed on the secure port (default 10259) at &lt;code&gt;/metrics/resources&lt;/code&gt;. It is a different endpoint from &lt;code&gt;/metrics&lt;/code&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Summary&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;kube_pod_resource_request{resource="memory", unit="bytes"}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The sum of requests per Pod. &lt;strong&gt;No series is emitted for a Pod whose value is 0&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;kube_pod_resource_limit{resource="memory", unit="bytes"}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The sum of limits per Pod&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

</description>
      <category>kubernetes</category>
      <category>linux</category>
      <category>cgroups</category>
      <category>prometheus</category>
    </item>
  </channel>
</rss>
