This is an English translation of an article I originally published in Japanese on Qiita.
TL;DR
Scope: Kubernetes v1.37, Linux 6.12, cgroup v2, containerd. Exact versions are listed in Components and versions.
"Memory usage" in Kubernetes is several different numbers, and OOM kills and evictions are each triggered by a different one of them. That is why a container can be OOMKilled while kubectl top shows headroom, or a Pod can be evicted while node utilization looks low.
Terms used throughout (details): memcg OOM is an OOM kill for exceeding a configured limit (container, Pod or /kubepods.slice), and global OOM is an OOM kill caused by memory exhaustion of the whole node. Eviction is a separate mechanism, done by the kubelet.
| Event | Triggering value | Condition |
|---|---|---|
| memcg OOM (OOM kill from exceeding a limit) |
memory.current (container_memory_usage_bytes) of the container, Pod or /kubepods.slice
|
Reaches memory.max at the same level and reclaim cannot bring it back under (a container's memory.max is its memory limit) |
| global OOM (OOM kill from node memory exhaustion) | Whether the whole node's memory is exhausted (from the kernel's point of view) | There is no threshold. The victim is chosen by oom_score (QoS class and memory request) |
| Eviction trigger | The node's memory.available (= MemTotal − node working set) |
Capacity falls below the evictionHard / evictionSoft thresholds (kubelet default: memory.available<100Mi) |
| Choosing which Pod to evict | Difference between a Pod's working set and its memory request | Pods whose actual usage exceeds their request are preferred |
Key points:
-
kubectl topandcontainer_memory_working_set_bytesshow the working set, not RSS. It includes active page cache,shmemand kernel memory. -
memcg OOM compares
memory.currentagainst the limit, not the working set, and only kills when reclaim cannot bring usage back under the limit. -
Eviction is decided from the node-wide working set, not
MemAvailable, so node-exporter utilization can look healthy while the kubelet is close to its threshold. -
OOMKilleddoes not tell you the path. ASystemOOMevent on the Node means global OOM, but its absence proves nothing (events are kept for 1 hour by default). - The most effective fix for global OOM and eviction is to set the memory request to match actual usage. It does not help with memcg OOM, which needs a limit review.
Where to go next:
- Why do the numbers disagree? Quick reference and Kubernetes memory usage and cgroup v2
- Something was just killed or evicted: Investigating OOM and eviction
- Prevent it: What to do about it, with PromQL for checking requests
- Specific questions: FAQ
Introduction
I get asked about Kubernetes memory usage again and again, and I kept giving the same explanation. The information you need is spread across the Kubernetes and Linux kernel documentation, and I could not find anything that connects the two, so I wrote it down myself.
There is no single definition of "memory usage"
The question "how many MiB of memory is this container using?" has no single correct answer. Unlike CPU utilization or disk usage, memory usage is not one value with one definition. The number changes depending on what you want to measure and where you measure it from.
The reason is that Linux cannot assign every byte of memory to exactly one owner.
-
Who owns the page cache? When you read a file, the kernel keeps its contents in memory. The application did not request this with
malloc; the kernel allocated it on its own for performance. And when memory runs short, the kernel silently throws it away. Whether it counts as "memory the application is using" depends on your purpose. - Who owns shared memory? Shared library pages, tmpfs and shared memory are referenced by several processes through the same physical pages. If you simply add up per-process numbers, the same page is counted many times, and the total can even exceed physical memory.
-
When does "freed" memory actually go down? Even if an application calls
free, the OS-visible usage does not drop unless the allocator or language runtime returns the memory to the OS. If it does return it, usage drops. The value depends on which layer you look at. - Some memory is reclaimable and some is not. Both count as "in use", but under memory pressure the kernel can take some of it back (page cache) and cannot take back the rest (anonymous pages). You cannot ignore this distinction when estimating free memory.
In other words, memory usage is a value that only becomes defined once you fix the measurement method and the purpose. "The pages a process holds in physical memory", "the memory only that process uses" and "the memory that cannot be freed under pressure" are all legitimate definitions, and each gives a different number.
Kubernetes runs on its own definitions
On top of that, Kubernetes components use their own definitions for management purposes. If you read a value assuming it is the same thing you normally call "memory usage", you will misinterpret it.
The container memory usage you see with kubectl top pod or container_memory_working_set_bytes is not the amount of memory the application has allocated (RSS). Kubernetes treats the working set as container usage, and it includes RSS plus part of the page cache and memory used by the kernel.
container_memory_working_set_bytes # the value kubectl top pod shows
≒ container_memory_rss # memory the application allocated (anonymous pages)
+ active page cache # only inactive pages are subtracted
+ tmpfs and shared memory (shmem) # included in full, active or inactive
+ kernel memory, etc.
So the working set is larger than what the application actually consumes. The tricky part is that OOM kills and evictions happen because a value other than the one you are looking at reached its limit. For an OOM kill caused by exceeding a limit (called memcg OOM in this article), the kernel compares the container's memory.current against the limit, which is not the working set. And memory.current merely reaching the limit does not cause an OOM; the kill happens when reclaim fails to bring usage back under the limit. Eviction is triggered by free memory computed from the node-wide working set, so it looks at a different scope than the per-container values.
If you treat these as the same thing, you cannot explain situations that look contradictory:
- The working set has not reached the memory limit, yet the container was
OOMKilled - Node memory utilization computed from
node_memory_MemAvailable_bytesis low, yet a Pod was evicted - The application does not leak memory, yet the working set keeps growing
- The RSS the application reports does not match
container_memory_rss
None of these are metric bugs. They are the result of each component reporting a value with a definition that suits its own purpose. Conversely, once you know which value has which definition, the apparent contradictions can be explained.
This article covers three things in order:
- Kubernetes memory usage and cgroup v2: what each value that appears as "memory usage" measures, which cgroup v2 value it corresponds to, and which value triggers eviction
-
How to investigate OOM kills and evictions: there are two OOM kill paths, and how to tell which one happened when you see
OOMKilledor an eviction - What to do about it: what to do once you know the cause. Right-sizing the memory request is the most effective way to reduce OOM kills and evictions together, so the article focuses on that, including PromQL for checking it
How this article names the two OOM kill paths
There are two paths by which a Pod's process gets killed for memory reasons, and they differ in both cause and blast radius. If you call both "OOM" you cannot tell them apart, so this article uses the following two names throughout.
| Name in this article | Purpose | Triggering cause | What gets killed |
|---|---|---|---|
| memcg OOM | Enforcing a configured limit | A container, Pod or /kubepods.slice exceeded its memory.max (= memory limit)
|
Only processes inside the cgroup that hit its limit |
| global OOM | Keeping the whole system alive (protecting the node) | Memory exhaustion of the whole node, regardless of any container's limit | Any process on the node |
Put simply, memcg OOM is caused by "your own limit", and global OOM is caused by "the node's free memory". Many cases of OOMKilled despite headroom under the limit are the latter.
The difference in purpose is written in the kernel documentation. Global OOM exists to keep the rest of the system alive; Concepts overview - OOM killer says "In order to save the rest of the system, it invokes the OOM killer". The description of memory.max, on the other hand, is "This is the main mechanism to limit memory usage of a cgroup" (Memory Interface Files), with no mention of protecting other cgroups. cgroup v2 treats limits (memory.max / memory.high) and protection (memory.min / memory.low) as separate models ("implements both limit and protection models"), and memcg OOM belongs to the former.
So memcg OOM can happen even when the node has plenty of free memory. It happens solely because a configured limit was exceeded, regardless of actual pressure. Global OOM, conversely, happens only when node memory is genuinely exhausted, regardless of any limit.
The one exception is the /kubepods.slice hierarchy. Its memory.max is set to Capacity - kube-reserved - system-reserved (see below), which effectively reserves memory for non-Pod processes such as the kubelet and containerd. That is not the purpose of the kernel's memcg OOM; it is Kubernetes using the mechanism that way.
Note: These two names are not official Kubernetes terms. Kubernetes only exposes the termination reason
OOMKilled, and it does not distinguish the paths (the reason is explained in Why OOMKilled cannot distinguish the paths).Both are terms the kernel itself uses, and this article adopts them as they are. The two are used side by side in a kernel comment (mm/oom_kill.c,
pagefault_out_of_memory()):/* * The pagefault handler calls here because some allocation has failed. We have * to take care of the memcg OOM here because this is the only safe context without * any locks held but let the oom killer triggered from the allocation context care * about the global OOM. */The actual distinction is whether the
memcgfield ofstruct oom_controlis NULL. If it holds a value, the victim is chosen from within that cgroup; if NULL, from the whole node (oom.h).is_memcg_oom()just checks this field./* Memory cgroup in which oom is invoked, or NULL for global oom */ struct mem_cgroup *memcg;
memcgstands for memory cgroup (memory control group), a short form ofstruct mem_cgroupin the code above. Variable and function names in the source (is_memcg_oom(),try_charge_memcg()) use the same spelling, and the kernel documentation also writes "Memory cgroup (memcg)" (Documentation/mm/hmm.rst).Note that the term
global OOMappears only in source comments and OOM logs; it is not in the kernel documentation.memcg OOMdoes appear as "memcg OOM killer" inDocumentation/admin-guide/cgroup-v1/memory.rst.This is how they appear in the kernel log:
Name in this article How it appears in the kernel log memcg OOM constraint=CONSTRAINT_MEMCG,oom_memcg=<cgroup path>,Memory cgroup out of memory: Killed process ...global OOM constraint=CONSTRAINT_NONE,,global_oom,Out of memory: Killed process ...The prefix of the
Killed processline directly indicates the path (oom_kill.c):oom_kill_process(oc, !is_memcg_oom(oc) ? "Out of memory" : "Memory cgroup out of memory");
Also note that eviction is a separate mechanism from an OOM kill. It is the kubelet, not the kernel, that removes the Pod, and it is observed as Evicted. It is covered below as well.
Quick reference: contradictions explained and a metric map
This section expands the TL;DR: it explains the apparent contradictions from the introduction and maps each metric to the cgroup v2 value behind it.
With that, the seemingly contradictory situations from the introduction can be explained as follows:
-
The working set has not reached the limit, yet
OOMKilled: memcg OOM comparesmemory.current, not the working set, against the limit (and merely reaching the limit does not trigger it; the kill happens when reclaim fails). Also, even if your container is under its limit, it can be killed in a node-wide memory exhaustion (global OOM). See My working set is below the limit, but the container was OOMKilled. -
Node utilization computed from
MemAvailableis low, yet a Pod was evicted:MemAvailablecounts part of the page cache as free, but the kubelet is working-set based and treats all active page cache as "in use". See Node usage computed from MemAvailable is low, but Pods are evicted. - The application does not leak memory, yet the working set keeps growing: the working set includes active page cache. Page cache is not reclaimed until memory gets tight, so workloads with heavy file I/O or logging look like they keep growing. See Memory working set keeps growing.
-
The RSS the application reports does not match
container_memory_rss: the former is per-process RSS (including file mappings), and the latter is the anonymous pages charged to the container's cgroup. See The RSS my application reports differs from the RSS metric.
The mapping between these values and the metrics you see in kubectl top and dashboards is as follows. Even among "memory usage" metrics, each one measures something different and corresponds to a different cgroup v2 field.
| Scope | Metric | Emitted by | Backing value in cgroup v2 / /proc
|
Page cache handling | Main use |
|---|---|---|---|---|---|
| Container | container_memory_working_set_bytes |
cAdvisor (embedded in kubelet) | memory.current − inactive_file |
Active pages count as in use |
kubectl top pod, usage visualization |
| Container | container_memory_rss |
cAdvisor (embedded in kubelet) |
anon in memory.stat
|
Not included (anonymous pages only) | Understanding what the app allocated |
| Container | container_memory_usage_bytes |
cAdvisor (embedded in kubelet) | memory.current |
Everything included | What memcg OOM is evaluated against |
| Node | container_memory_working_set_bytes{id="/"} |
cAdvisor (embedded in kubelet) |
anon + file − inactive_file in the root cgroup's memory.stat (see below) |
Active pages count as in use |
Eviction decisions, kubectl top node
|
| Node | node_memory_MemAvailable_bytes |
node-exporter |
MemAvailable in /proc/meminfo
|
Reclaimable pages count as free | Visualizing node memory utilization |
| Node | kube_node_status_capacity{resource="memory"} |
kube-state-metrics | (status.capacity of the Node object) |
(capacity, not usage) | Denominator when converting to utilization |
Looking at the same node, utilization computed from node_memory_MemAvailable_bytes does not match the utilization the kubelet uses for eviction decisions. With workloads that do a lot of file I/O or logging, the difference can reach several GiB.
Different emitters also mean different labels on the metrics. cAdvisor metrics carry no label identifying the node, so when you compute per-node utilization you need to align labels with kubelet_node_name, as shown later.
Finally, the conclusions for diagnosis and remedy:
-
To tell which OOM path occurred, check whether the Node has a
SystemOOMevent. If it does, it is global OOM, but its absence is not evidence of memcg OOM (neitherOOMKillednor metrics distinguish the paths). See Telling which path occurred. -
The most effective remedy is to set the memory request to match actual usage. It affects placement, eviction candidate selection and
oom_score_adjat the same time. PromQL for checking is collected in What to do about it.
Kubernetes memory usage and cgroup v2
This section sorts out which cgroup v2 value each metric corresponds to, and which value triggers OOM kills and evictions.
Memory breakdown in cgroup v2
Container memory usage is managed in a per-container cgroup. In cgroup v2, memory.current is the current usage and memory.stat is its breakdown.
# Example values in a container's cgroup (checked on the node)
$ cat /sys/fs/cgroup/<container cgroup>/memory.current
5368709120
$ cat /sys/fs/cgroup/<container cgroup>/memory.stat
anon 2147483648 # anonymous pages (heap, stack: memory not backed by a file)
file 3087007744 # page cache (including tmpfs and shared memory)
shmem 268435456 # the part of file that is tmpfs or shared memory
...
inactive_anon 134217728 # anon LRU: not accessed for a while
active_anon 2013265920 # anon LRU: recently accessed
inactive_file 1073741824 # file LRU: not accessed for a while
active_file 1744830464 # file LRU: recently accessed
unevictable 268435456 # pages that cannot be reclaimed (here, the tmpfs part)
...
The following relationships hold in this example. Note that inactive_file + active_file does not equal file; it equals file - shmem. The next section explains why.
anon + file + kernel memory = 5368709120 (memory.current)
inactive_file + active_file = 2818572288 = file - shmem
inactive_anon + active_anon = 2147483648 = anon (shmem is not in here)
unevictable = 268435456 = shmem (tmpfs without swap)
The main fields of memory.stat are:
anon (anonymous pages)
- Memory not backed by a file. Regions the application allocated with
mallocand similar belong here. - Unlike page cache, the kernel cannot reclaim it independently of the process. With swap it can write pages out and free up space, but on nodes that do not use swap it is memory the kernel cannot reclaim under pressure. When swap is unavailable, the kernel sets
scan_balance = SCAN_FILEand excludes the anon LRU from scanning (get_scan_count()in mm/vmscan.c). - On the other hand, this value drops if the process itself returns memory to the OS with
munmapormadvise(MADV_DONTNEED). Note that callingfreedoes not reduce it unless the allocator or language runtime returns the memory to the OS.
file (page cache)
- Contents of files held in memory by the kernel during reads and writes. The application did not allocate it explicitly.
- When memory runs short the kernel discards (reclaims) it automatically, so it is basically reclaimable memory.
- Cache of files on disk is split into
active_file(recently accessed) andinactive_file(not accessed for a while). Reclaim takesinactive_filefirst. -
active_file + inactive_filedoes not equalfile. These are totals of the LRU lists, and as described belowshmemis on the anon LRU rather than the file LRU. Besidesshmem, pages that fall off the LRU, such asmlocked pages, also show up as a difference. -
It includes not only the cache of files on disk but also
shmem(tmpfs and shared memory), because Linux treats these as page cache too.shmemis written out to swap rather than to a file, so on nodes without swap it cannot be reclaimed. The breakdown is visible asshmeminmemory.stat.
Warning:
shmemis counted infile, but it is not on the file LRU. The kernel classifies swap-backed pages as anon, and tmpfs pages are among them.folio_is_file_lru()in mm_inline.h is defined as "0 if@foliois a normal anonymous folio, a tmpfs folio or otherwise ram or swap backed folio".The counters they are charged to are different, though. A tmpfs page is added to both
NR_FILE_PAGES(=file) andNR_SHMEM(=shmem) (mm/shmem.c):__lruvec_stat_mod_folio(folio, NR_FILE_PAGES, nr); __lruvec_stat_mod_folio(folio, NR_SHMEM, nr);Furthermore, for a Kubernetes
emptyDirwithmedium: Memory, the kubelet mounts tmpfs with thenoswapoption (generateTmpfsMountOptions()in empty_dir.go). The kernel makes the inode of anoswaptmpfs unevictable (shmem_get_inode()in mm/shmem.c), so these pages sit on the unevictable LRU, not the anon LRU.if (sbinfo->noswap) mapping_set_unevictable(inode->i_mapping);So
shmemis accounted as follows:
mount LRU Where it shows up in memory.statemptyDirwithmedium: Memory(noswap)unevictable unevictableAny other tmpfs / shared memory anon LRU inactive_anon/active_anonIn neither case does it enter
inactive_file/active_file. That matters in two ways later:
inactive_file + active_fileequalsfile - shmem, notfile- Only
inactive_fileis subtracted from the working set, soshmemstays in the working set in full
With workloads that write a lot of logs or read and write large files, the page cache can grow to several GiB even though the application did not intend that.
Note:
memory.currentincludes memory the kernel itself uses (slab, socket buffers and so on) in addition toanonandfile. So addinganonandfilealone does not givememory.current.
How cgroup v2 memory controls map to Pod settings
cgroup v2 memory control interface
cgroup v2 offers four levels of control over memory usage. memory.min and memory.low are protections ("do not reclaim up to here"), and memory.high and memory.max are limits ("do not allow use beyond here").
| File | Type | Behavior |
|---|---|---|
memory.min |
Protection (hard) | Memory up to this value is not reclaimed. If nothing else can be reclaimed and usage cannot fit, an OOM kill occurs |
memory.low |
Protection (soft) | Best-effort avoidance of reclaim up to this value. It is reclaimed if nothing else can be |
memory.high |
Limit (soft) | Exceeding it applies heavy reclaim pressure and throttles the cgroup's processes. No OOM kill |
memory.max |
Limit (hard) | Exceeding it triggers reclaim, and if usage still does not fit, the cgroup's OOM killer runs |
Other files that report state:
-
memory.current: current usage (whatcontainer_memory_usage_bytesis backed by) -
memory.stat: usage breakdown -
memory.events: counts oflow,high,max,oom,oom_killandoom_group_kill -
memory.swap.max: swap limit -
memory.pressure: PSI for waiting on memory allocation
For exact definitions of each file see the kernel's Memory Interface Files, and for how Kubernetes treats cgroup v2 see About cgroup v2.
How Pod settings are reflected
Pod settings are reflected into cgroup v2 as follows.
| Pod setting | Target in cgroup v2 |
|---|---|
resources.limits.memory |
memory.max |
resources.requests.memory |
Not reflected (with the default kubelet configuration in v1.37; you can reflect it by changing settings → KEP-2570) |
resources.limits.cpu |
cpu.max |
resources.requests.cpu |
cpu.weight |
cgroup v2 is a directory tree. Its top is the root cgroup (/sys/fs/cgroup), where any process that is in no child cgroup belongs, so it contains not only Pods but every process on the node, such as the kubelet and containerd. Kubernetes creates the hierarchy below it: "all Pods on the node → Pod → container".
# With the systemd cgroup driver (the kubeadm default; names differ with the cgroupfs driver, see below)
/sys/fs/cgroup ← root cgroup (id="/")
│ contains every process on the node
│ the value used for eviction decisions
│
├── system.slice/ kubelet, containerd, sshd, etc.
│ ├── kubelet.service (processes that are not Pods)
│ └── containerd.service
│
└── kubepods.slice/ ← all Pods on the node
│ (id="/kubepods.slice")
├── kubepods-pod<UID>.slice/ ← Pod (Guaranteed)
│ ├── cri-containerd-<ID>.scope ← container
│ └── cri-containerd-<pause ID>.scope ← pause container
│
├── kubepods-burstable.slice/ (intermediate slice per QoS class)
│ └── kubepods-burstable-pod<UID>.slice/ ← Pod (Burstable)
│ └── cri-containerd-<ID>.scope ← container
│
└── kubepods-besteffort.slice/
└── kubepods-besteffort-pod<UID>.slice/ ← Pod (BestEffort)
└── cri-containerd-<ID>.scope ← container
When filtering metrics, the root cgroup is id="/", all Pods on the node are id="/kubepods.slice", a Pod is container="", image="", pod!="", and a container is container!="", image!="". What the kubelet uses for eviction decisions is the working set of the topmost root cgroup (see below). That is a different level from memcg OOM caused by a container or Pod exceeding its limit.
Memory limits are set not only on containers but also on higher levels.
| Level | cgroup | Memory limit |
|---|---|---|
| Whole node (root cgroup) | / |
Not set (the root has no memory.max → see below) |
| All Pods on the node | /kubepods.slice |
Capacity - kube-reserved - system-reserved |
| Pod | Differs by QoS class (see below) | The Pod's effective limit (see below). Set only when every container has a limit |
| Container |
cri-containerd-<ID>.scope under the Pod's cgroup |
That container's limits.memory
|
A Pod's effective limit is not simply the sum of its regular containers' limits.memory. If there are init containers, their maximum is also considered, and spec.overhead is added if set. For the calculation see Resource sharing within containers and Pod Overhead.
So even if a container has not reached its own limit, an OOM kill can occur when the Pod as a whole or /kubepods.slice hits its limit.
Note: The paths above assume the kubelet uses the systemd cgroup driver. With the cgroupfs driver, paths look like
/kubepods/burstable/pod<UID>/<container ID>, with no.sliceor.scopeand nocri-containerd-prefix on the container.The kubelet's own default is
cgroupfs, but kubeadm setssystemdwhen the value is empty (defaults.go, kubelet.go). You can check which one is in use with:$ kubectl get --raw "/api/v1/nodes/<node name>/proxy/configz" | jq '.kubeletconfig.cgroupDriver' "systemd"The
idlabel of cAdvisor metrics also changes with the cgroup driver, so when filtering in PromQL, check the value in your own cluster with this query:group by (id) (container_memory_working_set_bytes{id=~"/kubepods|/kubepods.slice"})A Pod's cgroup path depends on its QoS class. As the tree above shows, Burstable and BestEffort go through an intermediate slice per QoS class.
In the
<UID>part, systemd escaping replaces the hyphens of the Pod UID with underscores (example:6983351e-4ede-...→pod6983351e_4ede_...). The UID of a static Pod (mirror Pod) is a hash with no hyphens, so it looks likepod1ff70b81885169da4e9e32c9530b77c9.slice.The QoS class and UID of a Pod can be read from the Pod resource, and once you have those two you can assemble the cgroup path.
$ kubectl get pod -n <namespace> <Pod name> -o jsonpath='{.status.qosClass}{"\n"}{.metadata.uid}{"\n"}' Burstable 6983351e-4ede-4611-8bab-b986614af18d # To list them $ kubectl get pods -n <namespace> -o custom-columns='NAME:.metadata.name,QOS:.status.qosClass'
Container metrics
The cAdvisor built into the kubelet exposes the cgroup values above as Prometheus metrics. They are served from the kubelet's /metrics/cadvisor endpoint, and you do not need to run a separate cAdvisor Pod or DaemonSet.
| Metric | Emitted by | Backing value in cgroup v2 | Description |
|---|---|---|---|
container_memory_usage_bytes |
cAdvisor (embedded in kubelet) | memory.current |
Usage including anonymous pages, page cache and kernel memory. The value memcg OOM is evaluated against |
container_memory_working_set_bytes |
cAdvisor (embedded in kubelet) | memory.current - inactive_file |
Memory considered "not immediately reclaimable". The value Kubernetes treats as usage |
container_memory_rss |
cAdvisor (embedded in kubelet) |
anon in memory.stat
|
Amount of anonymous pages. Unlike true RSS, it does not include mmaped file pages (see below) |
container_memory_cache |
cAdvisor (embedded in kubelet) |
file in memory.stat
|
Page cache size. Includes shmem (tmpfs and shared memory) |
container_memory_total_active_file_bytes |
cAdvisor (embedded in kubelet) |
active_file in memory.stat
|
Active page cache |
container_memory_total_inactive_file_bytes |
cAdvisor (embedded in kubelet) |
inactive_file in memory.stat
|
Inactive page cache |
container_memory_mapped_file |
cAdvisor (embedded in kubelet) |
file_mapped in memory.stat
|
Amount of mmaped files |
container_pressure_memory_waiting_seconds_total |
cAdvisor (embedded in kubelet) |
total of some in memory.pressure
|
Time some processes were waiting for memory (see below) |
container_pressure_memory_stalled_seconds_total |
cAdvisor (embedded in kubelet) |
total of full in memory.pressure
|
Time all processes were stalled waiting for memory (see below) |
container_oom_events_total |
kubelet (v1.37; the name starts with container_ but it is not from cAdvisor) |
(not a cgroup value) | Counted by parsing oom-kill: lines in the kernel log. When the container terminates its cgroup disappears and the series is lost (whether or not it is recreated) |
Note: The kubelet has endpoints such as
/metrics,/metrics/cadvisorand/metrics/resource(server.go). The last two relate to container and node usage./metrics/cadvisorexposes thecontainer_*metrics above, while/metrics/resourceexposes only CPU, memory and swap usage (resource_metrics.go). The latter hascontainer_memory_working_set_bytesand also per-Pod and per-node series (pod_memory_working_set_bytes,node_memory_working_set_bytes) (seepodMemoryUsageDescandnodeMemoryUsageDescin resource_metrics.go). When to use which for per-Pod usage is covered in Align memory request with actual usage.
The important thing here is the definition of the working set.
container_memory_working_set_bytes
= container_memory_usage_bytes - container_memory_total_inactive_file_bytes
In other words, only inactive page cache is subtracted from the working set, and active page cache stays in as "in use". For the implementation see cAdvisor's handler.go.
workingSet := ret.Memory.Usage
if v, ok := s.MemoryStats.Stats[inactiveFileKeyName]; ok {
ret.Memory.TotalInactiveFile = v
if workingSet < v {
workingSet = 0
} else {
workingSet -= v
}
}
ret.Memory.WorkingSet = workingSet
As a breakdown, it is roughly the sum of the following. RSS (container_memory_rss) is exactly anon in memory.stat under cgroup v2, that is, the amount of anonymous pages, so it is included in the working set.
container_memory_working_set_bytes
≒ container_memory_rss (anonymous pages; shmem not included)
+ shmem (tmpfs and shared memory; stays in full)
+ container_memory_total_active_file_bytes (active page cache)
+ unreclaimable pages, kernel memory, etc. (unevictable, slab, socket buffers, ...)
shmem is listed as a separate term because neither container_memory_rss (= anon = NR_ANON_MAPPED) nor active_file includes shmem. As described above, shmem is not on the file LRU, so it never appears in active_file / inactive_file, and since only inactive_file is subtracted from the working set, it remains in full. No container_* metric is exposed for shmem, so this term can only be checked from memory.stat on the node.
It is ≒ because, besides shmem, there are also no exposed metrics for the kernel memory and locked pages contained in memory.current. For an exact breakdown, read memory.stat on the node directly.
Caution: RSS in
container_memory_rssstands for Resident Set Size, but its scope differs from the real meaning. True RSS means "the pages a process holds in physical memory" and includesmmaped file pages.But
memory.statin cgroup v2 has norssfield, and cAdvisor exposesanonascontainer_memory_rss. The kernel documentation definesanonas "Amount of memory used in anonymous mappings such as brk(), sbrk(), and mmap(MAP_ANONYMOUS)", which does not include file mappings. So it does not match theRSScolumn ofpsorVmRSSin/proc/[pid]/status.Warning: A large
container_memory_working_set_bytesdoes not necessarily mean the application needs that much memory. It may just be accumulated page cache. To see how much memory the application has actually allocated, compare againstcontainer_memory_rss.
Node metrics
There are two families of metrics representing whole-node memory usage, with different emitters.
From cAdvisor (the kubelet's view)
cAdvisor exposes memory usage not only for containers but also for the root cgroup (id="/", the topmost cgroup that contains every process on the node, as described above) and for the cgroup of all Pods. The id of the all-Pods cgroup depends on the cgroup driver, so the queries below match both values with =~.
-
container_memory_working_set_bytes{id="/"}: working set of the whole node (Pods + other processes). But the root cgroup has nomemory.current, so the backing value differs from containers (see below) -
container_memory_working_set_bytes{id=~"/kubepods|/kubepods.slice"}: working set of all Pods on the node -
container_memory_total_active_file_bytes{id="/"}: active page cache of the whole node
Both the MEMORY(bytes) column of kubectl top node and the kubelet's eviction decision described below use this working set.
From node-exporter (/proc/meminfo as is)
node-exporter turns the contents of /proc/meminfo directly into metrics.
-
node_memory_MemTotal_bytes: total physical memory on the node -
node_memory_MemAvailable_bytes: the kernel's estimate of "how much can be allocated when new memory is requested" -
node_memory_Cached_bytes: page cache size. As the kernel documentation says, "In-memory cache for files read from the disk (the pagecache) as well as tmpfs & shmem", it includes tmpfs and shared memory. The breakdown is innode_memory_Shmem_bytes
node_memory_MemAvailable_bytes counts reclaimable page cache as free. Computing node memory utilization as 1 - MemAvailable / MemTotal is also in node-exporter's official mixin as the NodeMemoryHighUtilization alert (alerts.libsonnet). Its expr uses variables for the selector and threshold; stripped of those it is 100 - (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes * 100) > <threshold>. However, as shown below, this value does not match the utilization the kubelet uses for eviction decisions.
Values the kubelet uses for eviction decisions
The kubelet checks the node's free memory (memory.available) every 10 seconds and starts evicting Pods when it falls below the threshold. This memory.available is not computed from MemAvailable in /proc/meminfo; it is computed from the root cgroup's working set.
memory.available = node_memory_MemTotal_bytes - container_memory_working_set_bytes{id="/"}
For the implementation see Kubernetes' helper.go and helpers_others.go. The 10-second interval is defined as evictionMonitoringPeriod.
Warning: The root cgroup working set used for eviction decisions (and by
kubectl top node) is not the same thing as a container's working set. In cgroup v2,memory.currenthas theCFTYPE_NOT_ON_ROOTflag, so the root cgroup has nomemory.current(memory_files[]in mm/memcontrol.c).{ .name = "current", .flags = CFTYPE_NOT_ON_ROOT, .read_u64 = memory_current_read, },
memory.stat, on the other hand, has no such flag, so it can be read even on the root. The cgroup library cAdvisor uses therefore substitutesanon + filefrommemory.statas the usage when readingmemory.currentfails withENOENT, for the root only (rootStatsFromMeminfo()in opencontainers/cgroups).// sum `anon` + `file` to report the same value as `usage_in_bytes` in v1. stats.MemoryStats.Usage.Usage = stats.MemoryStats.Stats["anon"] + stats.MemoryStats.Stats["file"]So for the
id="/"series:container_memory_usage_bytes{id="/"} = anon + file container_memory_working_set_bytes{id="/"} = anon + file - inactive_fileUnlike
memory.current, this includes no kernel memory at all, such as slab,pagetables,kernel_stack,percpuandsock. This is mainly why the node working set tends to come out smaller than usage derived from/proc/meminfo. For container and Pod cgroupsmemory.currentcan be read directly, so this substitution does not occur.
Because the working set includes active page cache, the kubelet treats active page cache as "memory in use".
The default threshold is only the hard eviction memory.available<100Mi, as in DefaultEvictionHard in defaults_linux.go; no soft eviction is configured by default.
In practice, though, I think it is better to override this value explicitly. In the v1.37 cluster I used for verification I also pass the following flags to the kubelet to set both hard and soft thresholds.
--eviction-hard=memory.available<5%,nodefs.available<5%,pid.available<5%
--eviction-soft=memory.available<10%,nodefs.available<10%,pid.available<10%
--eviction-soft-grace-period=memory.available=2m,nodefs.available=5s,pid.available=1m
The default 100Mi is a fixed value, so it becomes relatively thinner on nodes with more memory. If you specify percentages, the threshold follows the node size. Placing soft eviction ahead of hard (10 %) and giving it a grace period with --eviction-soft-grace-period avoids evicting Pods on momentary spikes while still evicting them when the pressure persists.
You can check your cluster's settings with the following query.
$ kubectl get --raw "/api/v1/nodes/<node name>/proxy/configz" | jq '.kubeletconfig | {evictionHard, evictionSoft, evictionSoftGracePeriod}'
{
"evictionHard": {
"memory.available": "5%",
"nodefs.available": "5%",
"pid.available": "5%"
},
"evictionSoft": {
"memory.available": "10%",
"nodefs.available": "10%",
"pid.available": "10%"
},
"evictionSoftGracePeriod": {
"memory.available": "2m",
"nodefs.available": "5s",
"pid.available": "1m"
}
}
Note: The denominator is Capacity (
node.status.capacity.memory,kube_node_status_capacity{resource="memory"}), not Allocatable. These are the same value asnode_memory_MemTotal_bytes. The kubelet uses cAdvisor's value as the machine's memory capacity, and cAdvisor reads it fromMemTotalin/proc/meminfo(raw/handler.go). A threshold likememory.available<5%is also evaluated as a percentage of this Capacity.Allocatable is Capacity minus the system reservation and the hard eviction threshold, and is the upper bound the scheduler uses when assigning Pod requests. It is not used to decide whether eviction happens. See Reserve Compute Resources for System Daemons for details.
Why node-exporter and kubelet usage disagree
Putting the above together, there are two views of the same node.
| View | Calculation | Emitted by | Active page cache |
|---|---|---|---|
| Node memory utilization from node-exporter | 1 - MemAvailable / MemTotal |
node-exporter (/proc/meminfo) |
The part estimated reclaimable counts as free |
| Kubelet's eviction decision | node working set / MemTotal |
cAdvisor (embedded in kubelet; root cgroup anon + file - inactive_file) |
Counted entirely as in use |
MemAvailable is the kernel's estimate of "how much can be newly allocated", and it adds part of the page cache and the reclaimable slab to free memory. It does not count the whole page cache as free. The implementation is si_mem_available() in mm/show_mem.c.
/*
* Estimate the amount of memory available for userspace allocations,
* without causing swapping or OOM.
*/
available = global_zone_page_state(NR_FREE_PAGES) - totalreserve_pages;
/*
* Not all the page cache can be freed, otherwise the system will
* start swapping or thrashing. Assume at least half of the page
* cache, or the low watermark worth of cache, needs to stay.
*/
pagecache = global_node_page_state(NR_ACTIVE_FILE) +
global_node_page_state(NR_INACTIVE_FILE);
pagecache -= min(pagecache / 2, wmark_low);
available += pagecache;
/*
* Part of the reclaimable slab and other kernel memory consists of
* items that are in use, and cannot be freed. Cap this estimate at the
* low watermark.
*/
reclaimable = global_node_page_state_pages(NR_SLAB_RECLAIMABLE_B) +
global_node_page_state(NR_KERNEL_MISC_RECLAIMABLE);
reclaimable -= min(reclaimable / 2, wmark_low);
available += reclaimable;
Three things can be read from this.
- What is subtracted is
min(half, watermark_low), so at least half of the page cache and the reclaimable slab is counted as free - From the free pages,
totalreserve_pagesis subtracted, notwatermark_low - The page cache term is the file LRU (
NR_ACTIVE_FILE + NR_INACTIVE_FILE), soshmem, which sits on the anon LRU, is never counted as free. On nodes that use tmpfs heavily,MemAvailablealso comes out small
The important point is that the two values always diverge, but which one is larger depends on the environment. There are four factors behind the divergence, and they push in opposite directions.
| Factor | Which utilization comes out higher |
|---|---|
The working set counts active_file entirely as "in use" |
The kubelet's utilization is higher |
The working set removes inactive_file entirely as "free" |
node-exporter's utilization is higher |
MemAvailable counts only about half of the page cache as "free" (shmem is never counted) |
node-exporter's utilization is higher |
The root's usage is computed as anon + file, so it includes no kernel memory such as slab or pagetables |
node-exporter's utilization is higher |
Three of the four factors push node-exporter's utilization higher, and only active_file pushes the kubelet's utilization higher. So the kubelet's utilization comes out higher only when active_file has grown to several GiB and outweighs the other three; otherwise node-exporter's utilization is higher.
The kubelet side is higher on nodes running workloads where active_file builds up easily, for example:
- Batch jobs and ETL that read and write large files
- Services that emit a lot of logs (stdout is written to files on the node by the container runtime, so it becomes page cache)
- Middleware designed around the page cache (Kafka, Elasticsearch, etc.)
- Log collection agents that keep reading log files on the node
What matters is the amount of file access, not the language or runtime. For example, a JVM heap is anonymous pages, so it raises RSS, not active_file.
On such a node it can happen that utilization computed from MemAvailable looks like 33 % while the kubelet's side has reached 53 %. If you look at the MemAvailable-based utilization and conclude there is plenty of room, the kubelet may in fact be approaching its threshold and evictions will occur.
On the other hand, on a node where page cache has not built up much:
MemTotal 3.820 GiB
MemAvailable 2.692 GiB → 1 - 2.692/3.820 = 29.6 % (node-exporter side)
usage(/) 2.557 GiB (= anon + file of the root)
inactive_file(/) 1.622 GiB
active_file(/) 0.227 GiB
Working Set(/) 0.935 GiB → 0.935/3.820 = 24.5 % (kubelet side)
active_file is only 0.23 GiB while inactive_file is 1.6 GiB, so the working set removes it all as "free", and the kubelet side comes out about 5 points lower. The kubelet side is not always higher.
Either way, the conclusion is that you cannot judge whether there is headroom from only one of the values. active_file is an auxiliary indicator for confirming the cause of the divergence, and its value is not itself the difference. When investigating node memory pressure, also check the kubelet-view utilization with the following PromQL.
# Node memory utilization from the kubelet's view (the value used in eviction decisions)
max by (node) (
container_memory_working_set_bytes{id="/"}
* on(instance) group_left(node) kubelet_node_name
)
/ max by (node) (kube_node_status_capacity{resource="memory"})
# Amount of active page cache, the main cause of the divergence
max by (node) (
container_memory_total_active_file_bytes{id="/"}
* on(instance) group_left(node) kubelet_node_name
)
Note: The label alignment is there because no cAdvisor
container_*metric carries any label identifying the node. Looking directly at the kubelet's/metrics/cadvisor, the root cgroup series looks like this:container_memory_working_set_bytes{container="",id="/",image="",name="",namespace="",pod=""} 8.92522496e+08The only thing that tells you the node is the
instancelabel Prometheus adds at scrape time, and labels such askubernetes_io_hostnameare added by your Prometheus relabel configuration. Because the name depends on the relabel configuration (or may not exist at all), referring to it withlabel_replaceis environment specific.So the queries above use
kubelet_node_name, which the kubelet exposes on/metrics.kubelet_node_name{node="memqos-control-plane"} 1This is a gauge whose value is always
1, and the node name is in thenodelabel. The value itself is meaningless; it is the info metric pattern, which carries the information in labels. The kubelet's help text says the same.// NodeName is a Gauge that tracks the node's name. The count is always 1. NodeName = metrics.NewGaugeVec( &metrics.GaugeOpts{ Subsystem: KubeletSubsystem, Name: NodeNameKey, Help: "The node's name. The count is always 1.",In OpenMetrics,
Infois defined as an independent metric type, with a mandatory_infosuffix and a sample value that is always1(OpenMetrics spec - Suffixes).Info metrics are used to expose textual information which SHOULD NOT change during process lifetime.
kubelet_node_namehas no_infosuffix and is exposed as a gauge, so strictly speaking it is not an OpenMetricsInfotype. But "value always 1, information in labels" is the same usage, so this article treats it as an info-metric pattern. Similar examples are kube-state-metrics'kube_node_infoand node-exporter'snode_uname_info. Both endpoints are scraped from the same target, soinstancematches, and* on(instance) group_left(node)brings in thenodelabel. And the resultingnodelabel has the same name as kube-state-metrics'nodelabel, so they join directly.If your environment adds its own node label by relabeling, you can use
label_replaceas follows (replacekubernetes_io_hostnamewith the value in your environment).max by (node) ( label_replace( container_memory_working_set_bytes{id="/"}, "node", "$1", "kubernetes_io_hostname", "(.*)" ) )Note: Eviction candidates are not limited to Pods in particular namespaces. When node memory gets tight, Pods in every namespace on that node are candidates. Which Pod is chosen is decided in this order: (1) whether usage exceeds the memory request, (2) Pod priority, (3) how far above the request it is. See Pod selection for kubelet eviction for details.
However, critical pods are excluded from selection and are never evicted at all (
evictPod()in eviction_manager.go). "Critical pod" is a term in the official Kubernetes documentation and refers to a Pod whose PriorityClass issystem-cluster-criticalorsystem-node-critical(Guaranteed Scheduling For Critical Add-On Pods says "To mark a Pod as critical, set priorityClassName for that Pod tosystem-cluster-criticalorsystem-node-critical.").if kubelettypes.IsCriticalPod(pod) { logger.Error(nil, "Eviction manager: cannot evict a critical pod", "pod", klog.KObj(pod)) return false }The kubelet's check is broader than this definition:
IsCriticalPod()is also true for static Pods and mirror Pods. The remaining condition is thatpod.Spec.Priorityis at leastSystemCriticalPriority(=2 × 1000000000), and since the upper limit of user-definable priority isHighestUserDefinablePriority(=1000000000), the only PriorityClasses that match are the two above. Pods with these two are not evicted, but in exchange they may not be protected in a global OOM. See Which process gets killed for details.
Investigating OOM and eviction
This section covers which values to check and in what order when OOMKilled or an eviction actually occurs.
The two OOM kill paths
There are two paths by which a container's process gets OOM killed. Both are SIGKILLs from the kernel's OOM killer, but the conditions under which they occur and the scope they affect differ. As described earlier, memcg OOM enforces a configured limit and global OOM protects the node, so their purposes differ at the root.
| Path | Condition | What gets killed | Node event |
|---|---|---|---|
| memcg OOM |
memory.max exceeded for a container, Pod or /kubepods.slice
|
Processes inside the cgroup that hit its limit | Not recorded |
| global OOM | Memory exhaustion of the whole node | Chosen from every process on the node by oom_score
|
SystemOOM is recorded if the kubelet detects it |
Even if your container is under its limit, it can be OOM killed because of another Pod on the same node. That is global OOM.
A Node event is recorded only by the kubelet. If SystemOOM exists it confirms global OOM, but events are kept for only one hour by default, so they become unfindable over time. So not finding a SystemOOM does not mean it was not a global OOM.
OOM kill from exceeding a cgroup limit (memcg OOM)
resources.limits.memory is set as the container cgroup's memory.max. When a cgroup's memory.current reaches memory.max, this is what happens. As described above, limits are set not only on containers but also on Pods and /kubepods.slice, and the same flow applies whichever level reaches its limit.
- The kernel tries to reclaim memory inside that cgroup. Page cache is basically reclaimed here.
- If usage still does not fit under
memory.maxafter reclaim, the cgroup's OOM killer SIGKILLs processes in that cgroup. - If the container's main process was killed,
lastState.terminated.reasonbecomesOOMKilledand the container is recreated according torestartPolicy. If only a child process was killed, the container keeps running, so neitherOOMKillednor a restart is observed.
The implementation is try_charge_memcg() in mm/memcontrol.c. It is a goto loop returning to the retry: label, and proceeds in this order.
| Step | Location | What happens |
|---|---|---|
| 1. Try to charge | L2177-L2178 | If page_counter_try_charge() succeeds, the allocation completes and it returns |
| 2. Reclaim | L2211-L2216 | On failure it reclaims with try_to_free_mem_cgroup_pages(), and if mem_cgroup_margin() satisfies the request, goes back to 1 |
| 3. Count the retry | L2244-L2245 | If reclaim is not enough, it decrements nr_retries and goes back to 1. The initial value is MAX_RECLAIM_RETRIES (= 16; mm/internal.h) |
| 4. Invoke the OOM killer | L2259-L2264 | Only after the retries are used up does it call mem_cgroup_oom()
|
An OOM happens only after up to 16 reclaim-and-retry rounds. The check in step 2 looks like this in code.
nr_reclaimed = try_to_free_mem_cgroup_pages(mem_over_limit, nr_pages,
gfp_mask, reclaim_options, NULL);
psi_memstall_leave(&pflags);
if (mem_cgroup_margin(mem_over_limit) >= nr_pages)
goto retry;
The kernel documentation for memory.max says the same.
Memory usage hard limit. ... If a cgroup's memory usage reaches this limit and can't be reduced, the OOM killer is invoked in the cgroup.
Warning:
memory.currentreachingmemory.maxdoes not by itself cause an OOM. A cgroup sitting at its limit with page cache is a normal state. For the same reason, the working set reachingmemory.maxis not an OOM condition either. Onlyinactive_fileis subtracted from the working set, and reclaimable memory remains in the rest:
active_file: demoted to inactive under pressure and then reclaimedslab_reclaimable(dentry / inode caches): in memcg reclaim too,shrink_slab()is called aftershrink_lruvec()(shrink_node_memcgs()in mm/vmscan.c)So there is no threshold of "OOM once either value reaches the limit". The actual condition is when the charge fails and reclaim cannot make room. What cannot be reclaimed is mainly anonymous pages (without swap) and
shmem, so that is where to look when narrowing down.
That said, what the kernel compares against memory.max in a memcg OOM is the whole of memory.current, and it is not decided by any single metric. memory.current includes, besides anonymous pages, kernel memory, shmem (tmpfs and shared memory), and page cache that cannot be reclaimed, or cannot be reclaimed right away.
container_memory_rss is an indicator for determining whether anonymous pages are the main cause. If RSS is close to the limit you can conclude anonymous pages are the main cause, but RSS being below the limit is no evidence that an OOM will not occur. When the gap between RSS and the working set is large, check container_memory_usage_bytes (the value of memory.current) together with the breakdown in memory.stat on the node: anon, file, shmem, slab and so on.
OOM kill from node memory exhaustion (global OOM)
When the memory of the whole node is exhausted, the kernel's OOM killer evaluates every process on the node, picks the one with the highest oom_score, and SIGKILLs it. Whether an individual container exceeds its limit is irrelevant.
However, what acts first against node memory pressure is the kubelet's eviction. The kubelet evicts Pods to protect the node, and its selection puts "Pods whose usage exceeds their memory request" first (rankMemoryPressure() in helpers.go).
// rankMemoryPressure orders the input pods for eviction in response to memory pressure.
// It ranks by whether or not the pod's usage exceeds its requests, then by priority, and
// finally by memory usage above requests.
func rankMemoryPressure(pods []*v1.Pod, stats statsFunc) {
orderedBy(exceedMemoryRequests(stats), priority, memory(stats)).Sort(pods)
}
A Pod with neither limits nor requests always falls on the "exceeds the request" side, so it is chosen first. So if this kind of Pod makes the node tight and eviction is in time, it is observed as Evicted, not OOMKilled.
Global OOM happens first when eviction cannot reclaim memory as fast as memory grows. There are several reasons it cannot keep up.
| Factor | Value | Effect |
|---|---|---|
| Evaluation interval | 10 seconds (evictionMonitoringPeriod) |
Detection can be up to 10 seconds late |
| Pods evicted per evaluation | Only one ("we kill at most a single pod during each eviction interval" in eviction_manager.go) | Reclaim speed is limited to "one Pod per 10 seconds" |
| Waiting for cleanup | Up to 30 seconds (podCleanupTimeout) |
The next evaluation may be delayed further |
| Default threshold | memory.available<100Mi |
Extremely thin on large nodes |
How memory.available is computed |
From the root cgroup's anon + file
|
Contains no kernel memory, so it overestimates free memory |
The official Kubernetes documentation also says (Node out of memory behavior):
If the node experiences an out of memory (OOM) event prior to the kubelet being able to reclaim memory, the node depends on the oom_killer to respond.
It is more likely when applications that allocate a lot of memory right at startup, or in response to a sudden surge in requests, share a node.
Warning: Eviction protects the node; it does not protect individual Pods from OOM. Containers without
limitshave nomemory.maxbackstop at all, so eviction is merely the last line of defense.Moreover,
rankMemoryPressure()looks at Pod priority before usage size. If the offending Pod has high priority, low-priority Pods with small usage get evicted first, and global OOM can occur without the pressure being resolved. The proper countermeasure is to setlimitsandrequestsrather than relying on eviction.
Which process gets killed
The kernel kills the process with the highest oom_score. The kubelet sets oom_score_adj (a bonus added to oom_score) according to the container's QoS class, so the QoS class and the memory request setting directly become how likely a container is to be killed. You can check the QoS class with kubectl get pod <Pod name> -o jsonpath='{.status.qosClass}', and how it is determined is described in Pod Quality of Service Classes.
| Condition | oom_score_adj |
Likelihood of being killed |
|---|---|---|
system-node-critical (evaluated before the QoS class) |
-997 |
Least likely |
| Guaranteed | -997 |
Least likely |
| Burstable |
1000 - 1000 × memory request / node memory capacity (clamped to 3–999) |
The smaller the request, the more likely |
| BestEffort | 1000 |
Most likely |
The implementation is GetContainerOOMScoreAdjust() in Kubernetes' policy.go. system-node-critical is evaluated before the QoS class decision.
if types.IsNodeCriticalPod(pod) {
// Only node critical pod should be the last to get killed.
return guaranteedOOMScoreAdj
}
The Burstable clamp is implemented with a lower bound of 1000 + guaranteedOOMScoreAdj (= 3) and an upper bound of besteffortOOMScoreAdj - 1 (= 999). A Pod whose request is set smaller than its actual usage is more likely to be targeted in both eviction and global OOM.
Warning: PriorityClass works differently for eviction and for global OOM. It is a misreading to think that raising priority protects you from both.
PriorityClass Eviction global OOM ( oom_score_adj)User defined (up to 1000000000)The higher the value, the later it is evicted Determined by QoS class and request system-cluster-criticalExcluded (never evicted) Determined by QoS class and request (not protected) system-node-criticalExcluded (never evicted) Fixed at -997static / mirror Pod Excluded (never evicted) Determined by QoS class and request The three "excluded" rows in the Eviction column are the critical pods described earlier (Pods for which
IsCriticalPod()is true).What
oom_score_adjprotects is onlysystem-node-critical, for whichIsNodeCriticalPod()is true. If asystem-cluster-criticalPod is BestEffort, itsoom_score_adjis1000, and it is killed first in a global OOM. "I raised the priority but it was stillOOMKilled" comes from this path.Warning: Even in a global OOM, the Pod is not deleted as in eviction. What happens next depends on
restartPolicy.
AlwaysorOnFailure: the container is recreated on the same node. The Pod stays on the same node, so if the cause is not resolved, OOM kills will happen againNever: the container is not recreated and the Pod remainsFailed. If a controller such as a Job creates a successor Pod, it may be placed on a different nodeIf
OOMKilledrepeats even though there is headroom under the limit, suspect this path.Note: Because the OOM kill frees memory, the pressure may already be gone when the kubelet next evaluates free memory. In that case the node gets no
MemoryPressurecondition and nonode.kubernetes.io/memory-pressure:NoScheduletaint, and scheduling of new Pods does not stop. Conversely, if the pressure continues, a later evaluation sets the condition and taint, and eviction also occurs. Note that it is decided bymemory.availableat evaluation time, not by whether an OOM kill happened.
Telling which path occurred
Using Kubernetes information
Start by narrowing down with what you can check using kubectl.
| Where to look | What it tells you |
|---|---|
Last State in kubectl describe pod, or lastState.terminated in kubectl get pod -o yaml
|
That an OOM kill happened (reason: OOMKilled). It cannot distinguish the path
|
Events in kubectl describe node <node name>
|
If SystemOOM is recorded, it is a global OOM. If there is no record, that does not make it a memcg OOM; it just means you cannot tell |
There is no OOM-specific event on the Pod side. You only see things like BackOff accompanying the restart.
Check it as follows. SystemOOM is an event attached to the Node, and since Node is not a namespaced object, the event itself is recorded in the default namespace.
# 1. Find which node the OOMKilled Pod is on
$ kubectl get pod -n <namespace> <Pod name> -o wide
NAME READY STATUS RESTARTS AGE IP NODE
myapp-... 1/1 Running 5 (2m ago) 1h 10.0.0.1 node-1
# 2. Check whether that node's events contain SystemOOM
$ kubectl describe node node-1
...
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning SystemOOM 3m kubelet System OOM encountered, victim process: java, pid: 12345
# To list only events
$ kubectl get events -n default --field-selector reason=SystemOOM,involvedObject.name=node-1
If SystemOOM exists, a global OOM definitely happened on that node. However, the victim is not necessarily the container you are investigating, so also check that the time of SystemOOM corresponds to the container's termination time (lastState.terminated.finishedAt).
Not finding SystemOOM is not evidence of memcg OOM. The same state results from event TTL expiry, a kubelet restart, or a missed event.
The event retention period is set by kube-apiserver's --event-ttl, and the default is 1 hour (options.go). So a SystemOOM older than an hour cannot be seen with kubectl. If you forward events to an external logging platform, check there.
When you cannot tell, infer from the following:
- Whether
container_memory_usage_byteswas close to the limit at that time (if so, memcg OOM is likely) - Whether the kubelet-view node memory utilization was high at that time (if so, global OOM is likely)
Why OOMKilled cannot distinguish the paths
When a task exits with code 137, the container runtime (containerd) checks oom_kill in the cgroup's memory.events, and if it is 1 or more, it updates the termination reason to OOMKilled (events.go). The check is oomMetricsEventOccurred() in the same file.
switch v := taskMetricsAny.(type) {
case *cg1.Metrics:
return v.GetMemoryOomControl().GetOomKill() > 0, nil
case *cg2.Metrics:
return v.GetMemoryEvents().GetOomKill() > 0, nil
The definition of oom_kill in memory.events is "The number of processes belonging to this cgroup killed by any kind of OOM killer". So a process chosen and killed by the kernel in a global OOM is counted too, and both paths produce OOMKilled.
Why SystemOOM is recorded only for global OOM
SystemOOM is an event the kubelet records on the Node object. The kubelet parses OOM messages in the kernel log and records it only when the cgroup where the OOM happened is judged to be the root (/) (oom_watcher_linux.go).
for event := range outStream {
// Count every OOM kill per container to back the
// container_oom_events_total metric.
recordOOMKill(event.ContainerName)
if event.VictimContainerName == recordEventContainerName {
...
ow.recorder.WithLogger(logger).Eventf(ref, v1.EventTypeWarning, systemOOMEvent, "%s", eventMsg)
}
}
That this lines up exactly with global OOM comes from the combination of the kernel log format and the parser. For an OOM caused by a memcg limit, the kernel prints oom_memcg=<cgroup path>, whereas for a global OOM it prints ,global_oom instead (mem_cgroup_print_oom_context() in mm/memcontrol.c).
if (memcg) {
pr_cont(",oom_memcg=");
pr_cont_cgroup_path(memcg->css.cgroup);
} else
pr_cont(",global_oom");
The parser side only matches the format containing oom_memcg= (oomparser.go).
containerRegexp = regexp.MustCompile(`oom-kill:constraint=(.*),nodemask=(.*),cpuset=(.*),mems_allowed=(.*),oom_memcg=(.*),task_memcg=(.*),task=(.*),pid=(.*),uid=(.*)`)
So no cgroup can be extracted from a global OOM message, and the initial value / of OomInstance is left as is (StreamOoms in oomparser.go).
oomCurrentInstance := &OomInstance{
ContainerName: "/",
VictimContainerName: "/",
TimeOfDeath: msg.Timestamp,
}
In other words, the presence of SystemOOM is not "the result of directly determining the path" but "the result of whether the kernel log could be parsed". Still, in practice it is enough to understand that if it was recorded, you can conclude it was a global OOM, while its absence is no evidence at all.
How it appears in metrics
Metrics do not distinguish the path with a label. The kernel's __oom_kill_process() records the same event regardless of the path.
| Metric | Emitted by | memcg OOM | global OOM |
|---|---|---|---|
node_vmstat_oom_kill |
node-exporter (oom_kill in /proc/vmstat) |
Increases | Increases |
container_oom_events_total |
kubelet (result of kernel log parsing) | The series of the container the killed process belonged to increases | The id="/" series increases, not a container series |
oom_kill in cgroup v2 memory.events
|
(a cgroup file, not a metric; containerd reads it) | Increases | Increases |
Only container_oom_events_total behaves differently because, as described earlier, the parser cannot extract a cgroup in a global OOM and treats it as the root (/). This is only a result of whether the log could be parsed, so do not use it to determine the path; decide by the presence of SystemOOM.
Note: Although
container_oom_events_totalhascontainer_in its name, in v1.37 it is counted by the kubelet, not cAdvisor (recordOOMKill()in oom_counter.go). The meaning of the value does not change, but you have to look elsewhere when reading the code.
What to look at after narrowing down
- If you judge it to be a global OOM, the cause is memory exhaustion of the whole node. Compare the memory requests and actual usage of the Pods placed on the same node, and check whether any Pod has too small a request
-
If you judge it to be a memcg OOM, review that container's memory limit and its actual usage. If the container seems to have headroom under its limit, it may have reached the limit of the Pod as a whole or of
/kubepods.slice - If you cannot decide, keep both possibilities open and check both the container's limit and usage and the node's memory utilization
Investigation workflow
When investigating memory-related events, checking in the following order makes it easier to narrow down the cause.
1. When a Pod was OOMKilled
- Check whether
container_memory_usage_bytesis close to the limit. memcg OOM is evaluated against this value - Compare with
container_memory_rssto tell whether anonymous pages are the main cause. An OOM can occur even when RSS is low - If the gap between the working set or usage and RSS is large, page cache,
shmemor kernel memory may be pushing usage up - If
OOMKilledoccurs despite headroom under the limit, suspect global OOM. If the node's events have aSystemOOM, it is a global OOM. If there is none, it is not necessarily a memcg OOM - If you see no spike in the metrics, suspect a surge shorter than the scrape interval. See My working set is below the limit, but the container was OOMKilled for what to check
2. When a Pod was evicted
- If the message is
The node was low on resource: memory., node-wide memory pressure is the cause - Use the PromQL above to check the kubelet-view node memory utilization. The node-exporter-based utilization alone is not enough to judge
- Check whether a Pod with a small (or unset) memory request relative to actual usage is on the same node. Compare each Pod's
container_memory_working_set_bytes(cAdvisor) againstkube_pod_resource_request{resource="memory"}(kube-scheduler). Note that Pods with large requests make the scheduler keep placement density low, so they are unlikely to be the cause of pressure
3. To prevent recurrence
- Set a memory request that matches actual usage. It has the largest effect. PromQL for checking is collected in Align memory request with actual usage
- For containers where memcg OOM occurs, also review the limit
- Control placement with Pod anti-affinity or topology spread constraints so that Pods with large memory usage do not gather on the same node
Checking with PSI (Pressure Stall Information)
This is an auxiliary indicator for the case where things are slow although usage has not reached the limit. It shows how long processes were actually stalled waiting for memory reclaim.
You can obtain the time a container was stalled waiting for memory allocation as PSI. Even when usage has not reached the limit, you can detect being made to wait by memory reclaim. Note: These metrics are controlled by the The unit is seconds, so the result of When there is plenty of memory, neither increases, so To see memory pressure on the whole node, use the root cgroup series. Containers that are struggling although they are below their limit can be found by combining PSI with utilization. Here is an example alert. It catches The right threshold of For details on PSI see Understand Pressure Stall Information (PSI) Metrics.How to read PSI and PromQL
Metric
Emitted by
Description
container_pressure_memory_waiting_seconds_totalcAdvisor (embedded in kubelet)
Time some processes were waiting for memory (
total of some in memory.pressure)
container_pressure_memory_stalled_seconds_totalcAdvisor (embedded in kubelet)
Time all processes were stalled waiting for memory (
total of full in memory.pressure)KubeletPSI feature gate, which went GA in v1.36 and is enabled by default (it cannot be disabled). The kernel must support PSI, though, and the kubelet decides this by whether cpu.pressure exists at the cgroup root (cadvisor_linux.go). If the metrics do not appear, check whether the kernel was built with CONFIG_PSI=y and whether psi=0 is set as a kernel parameter.memory.pressure also has avg10 / avg60 / avg300, "the average over the last N seconds (%)", but only the cumulative value (total) is exposed as a Prometheus metric. So use rate() to see a ratio.rate() is "how many seconds of waiting per second", a ratio from 0 to 1. Multiply by 100 to get a percentage directly.
# Per Pod, the share of time (%) all processes were stalled waiting for memory
100 * max by (namespace, pod) (
rate(container_pressure_memory_stalled_seconds_total{container="", image="", pod!=""}[5m])
)
# Per Pod, the share of time (%) some processes were waiting for memory
100 * max by (namespace, pod) (
rate(container_pressure_memory_waiting_seconds_total{container="", image="", pod!=""}[5m])
)
rate() stays at 0. Under pressure, waiting (some) rises first, and if it gets worse stalled (full) follows. waiting counts even when only some processes waited, so to judge severity, use stalled. It is the time all processes in the container were stopped, so if it shows a value, throughput is directly affected.
# Share of time (%) the whole node was completely stalled waiting for memory
100 * max by (node) (
rate(container_pressure_memory_stalled_seconds_total{id="/"}[5m])
* on(instance) group_left(node) kubelet_node_name
)
container_spec_memory_limit_bytes is the container limit exposed by cAdvisor.
# Containers whose usage is under 80 % of the limit but that are stalled waiting for memory
(
100 * rate(container_pressure_memory_stalled_seconds_total{container!="", image!=""}[5m]) > 1
)
and
(
container_memory_usage_bytes{container!="", image!=""}
/ container_spec_memory_limit_bytes{container!="", image!=""} < 0.8
)
and keeps only series whose label sets match exactly on both sides. A container with no limit has container_spec_memory_limit_bytes of 0, so the division gives +Inf, which does not satisfy < 0.8, and it is excluded automatically.stalled that keeps occurring.
- alert: ContainerMemoryStalled
# Fires when the 5-minute average of fully stalled time is 5 % or more for 10 minutes
expr: |
100 * max by (namespace, pod) (
rate(container_pressure_memory_stalled_seconds_total{container="", image="", pod!=""}[5m])
) > 5
for: 10m
labels:
severity: warning
annotations:
summary: "A Pod is stalled waiting for memory allocation (add the namespace and pod labels with your own template)"
description: "It may be made to wait by reclaim even though it has not reached its memory limit. Check the gap between usage and RSS and the amount of page cache."
5 % depends on the environment. First look at the distribution of values without an alert to understand the normal level, then decide.
What to do about it
This is what to do once you know the cause. The most effective way to reduce OOM kills and evictions together is to right-size the memory request, so that is the focus.
Align memory request with actual usage
As we have seen, the most effective way to keep OOMs down is to set the memory request to match actual usage.
resources.requests.memory is not set in the cgroup, so it does not directly limit how much memory a container can use. It nevertheless matters because the request determines how likely OOM kills and evictions are, in these three ways:
- Placement (how likely both global OOM and eviction are): the scheduler decides placement by looking only at requests, not actual usage. A Pod with no request, or one smaller than actual usage, is placed while overestimating the node's free memory. As a result, Pods with large memory usage concentrate on the same node and the whole node becomes tight, inviting both global OOM and eviction
- Eviction candidate selection: the kubelet preferentially evicts Pods whose actual usage exceeds their memory request. If you set the request smaller than actual usage, that Pod is evicted more easily
-
oom_score_adj(the order of kills in a global OOM): as described earlier, the Burstableoom_score_adjis1000 - 1000 × memory request / node memory capacity. The smaller the request, the larger theoom_score_adj, and the more likely the Pod is chosen as a kill target in a global OOM. With no request it is BestEffort and1000, so it is killed first
"Whether the event happens" and "which Pod is chosen" are separate matters, so organized by effect:
| How the request acts | memcg OOM occurs | global OOM occurs | global OOM target selection | Eviction occurs | Eviction target selection |
|---|---|---|---|---|---|
| Placement (scheduler) | — | ○ | — | ○ | — |
oom_score_adj |
— | — | ○ | — | — |
| Eviction candidate selection | — | — | — | — | ○ |
Only placement affects whether an event occurs, and it affects both global OOM and eviction, because both start from node pressure. The effect is indirect, though: the request does not appear in the computation of memory.available (makeMemoryAvailableSignalObservation() mentioned earlier), so raising a request does not relax the threshold. The node just becomes less likely to get tight as a result of lower placement density.
The other two only affect whether your Pod is chosen after the pressure has built up. And the only thing the request does not affect is memcg OOM.
Conversely, just aligning the request with actual usage gives two effects at once: "the node is less likely to get tight" and "even when it does, your Pod is less likely to be chosen". Unlike increasing the limit, right-sizing the request does not change the total memory usable across the node, so it is also less likely to translate into cost.
Warning: Right-sizing the request does not help with memcg OOM (exceeding a container's limit). To reduce memcg OOM you need to review the limit. This section is about measures to reduce global OOM and eviction. Use Telling which path occurred to determine which one is happening.
PromQL for comparing requests and actual usage
You can check whether a request matches actual usage with PromQL. I included queries for the per-Pod ratio, for finding Pods with no request, for per-node comparison, and for deriving a recommended value from past usage. I also explain the pitfalls that come from where the metrics originate (such as double counting).
What to know before comparing A Pod's request is available as Which metric gives per-Pod memory usage kube-state-metrics converts Kubernetes API objects (spec and status) directly into metrics and does not handle actual resource usage at all. It provides configured values written in manifests, such as There are three ways to get actual per-Pod usage. The second and third are the same value. The kubelet gets If you do not scrape Warning: Do not omit For example, in a cluster running 100 Pods, the series count splits into these three kinds. With If you do In containerd environments the pause container's The queries from here on use the form that looks at the Pod's cgroup directly, which works in any environment. Compare request and actual usage for each Pod This is the ratio of actual usage to the request. A Pod above 1 has too small a request, and is more likely to be targeted in both eviction and global OOM. To see the difference in bytes, do the following. A positive value is "the amount by which the request is exceeded". Find Pods with no request As noted above, Pods with no request do not appear in Compare the sum of requests and the sum of actual usage per node For each node, compare "usage the scheduler knows about (sum of requests)" with "actual usage". A node where actual usage is larger is one where the scheduler overestimates free space, and the risk of global OOM and eviction is high. The actual-usage side comes from cAdvisor and has no Put these two side by side and look for nodes where the lower query (actual usage) exceeds the upper one (requests). Derive a recommended request from past usage Decide the request not from an instantaneous value but from usage over a period. Memory, unlike CPU, is a resource that cannot be reclaimed (incompressible), so it is safer to base it on a value near the peak, not the average. If you do not want to be pulled around by transient spikes, use a quantile. Be aware, though, that if you use this value as the request, the Pod will exceed its request at peak times and become an eviction candidate. Warning: The maximum obtained by this query misses spikes shorter than Prometheus' scrape interval. For applications that temporarily allocate a lot of memory right after startup, the real peak may be higher than this value. It is the same reason as in My working set is below the limit, but the container was OOMKilled. Leave some margin, or decide after understanding the application's startup behavior. Also, the working set includes active page cache. For containers with a lot of logging or file I/O, using this value as is for the request makes it too large. Compare with PromQL for checking and notes on filters
kube_pod_resource_request, which kube-scheduler exposes on its /metrics/resources endpoint. You compare it with actual usage (container_memory_working_set_bytes), but the two come from different sources, so dividing them as is gives the wrong result. Keep these five points in mind.
Point
Details
Scrape configuration is needed
kube_pod_resource_request is on kube-scheduler's secure port (default 10259) at /metrics/resources. It is a different endpoint from /metrics, so Prometheus needs explicit configuration
No series for Pods with no request
Because the collector does
if val.IsZero() { return }, Pods whose request is 0 or unset do not appear in the metric. The most problematic Pods drop out of the query, so find them separately with unless below
Aggregated per Pod
kube_pod_resource_request is the total for the whole Pod and has no container label. Actual usage must also be aligned to Pod level before comparing (how is shown below)
Duplicate scrapes
kube-scheduler runs multiple replicas in an HA setup, so the same Pod's series may be scraped several times with different
instance values. Summing them double counts, so normalize with max by (...) before aggregating
Terminated Pods are excluded
Pods that are Succeeded / Failed are not included
kube_pod_resource_request is a per-Pod value, so actual usage must also be aligned to Pod level. You may wonder "doesn't kube-state-metrics have a Pod memory usage metric?", but it does not.kube_pod_container_resource_requests / kube_pod_container_resource_limits, but usage is out of scope. Usage requires reading cgroups, which is the job of cAdvisor and the kubelet.
Method
Metric
Emitted by
Characteristics
Sum the containers
sum by (namespace, pod) (container_memory_working_set_bytes{...})cAdvisor (
/metrics/cadvisor)Works in any environment. Double counts if the filter is wrong (see below)
Look at the Pod's cgroup directly
container_memory_working_set_bytes{container="", image="", pod!=""}cAdvisor (
/metrics/cadvisor)No summing needed. Slightly larger than the container total because it includes the pause container and Pod-level cgroup charges.
image="" is required (see below)
The kubelet's per-Pod metric
pod_memory_working_set_bytes{namespace, pod}kubelet
/metrics/resource
The most straightforward, but it is on a different path from
/metrics/cadvisor, so you must add a scrape configuration to Prometheuspod_memory_working_set_bytes from the Pod-level cgroup, not as a sum of containers.
podUID := types.UID(podStats.PodRef.UID)
// Lookup the pod-level cgroup's CPU and memory stats
podInfo := getCadvisorPodInfoFromPodUID(podUID, allInfos)
if podInfo != nil {
cpu, memory := cadvisorInfoToCPUandMemoryStats(podInfo)
podStats.CPU = cpu
podStats.Memory = memory
pod_memory_working_set_bytes is a STABLE metric, so if Prometheus scrapes /metrics/resource in your environment, using it is the simplest.
# Actual per-Pod usage (when scraping the kubelet's /metrics/resource)
pod_memory_working_set_bytes
/metrics/resource, specify the Pod cgroup series directly. /metrics/resource is a different path from /metrics/cadvisor, so it is not collected unless you add it to the Prometheus configuration.
# Actual per-Pod usage (looking at the Pod's cgroup series directly)
max by (namespace, pod) (
container_memory_working_set_bytes{container="", image="", pod!=""}
)
container="" gets the Pod's cgroup because when the kubelet labels metrics it also attaches pod and namespace to the Pod's cgroup series. The code explains it as "Associate pod cgroup with pod so we have an accurate accounting of sandbox" (server.go).image="". If you filter with only container="", the series of the pause container (sandbox) also matches along with the Pod's cgroup. In containerd environments the pause container's container label is an empty string and its image label holds the pause image.
Selector
Series
What it is
{container="", pod!=""}200
Pod cgroups + pause containers
{container="", image="", pod!=""}100
Pod cgroups only
{container="", image!="", pod!=""}100
Pause containers only
max by (namespace, pod) the Pod's cgroup is larger, so the correct value comes back, but with sum by the pause part is added and it is double counted. For one Pod it looks like this:
Pod cgroup 666,116,096 ← correct value
Sum of real containers 665,882,624
pause container 225,280
sum{container="",pod!=""} 666,341,376 ← Pod cgroup + pause, double counted
sum by (namespace, pod) without a container filter, each container's share and the Pod cgroup's share are both added, so the value is almost doubled. When summing containers, narrow to real containers only, as follows.
# Actual per-Pod usage (when summing containers)
sum by (namespace, pod) (
container_memory_working_set_bytes{container!="", image!=""}
)
container is empty, so container!="" alone excludes the Pod cgroup, the root cgroup and pause. image!="" is added because series other than real containers have no image, and it is insurance for environments where container!="" alone leaves some series. With CRI-O the pause container appears as container="POD", so also add container!="POD" there.
# Ratio of actual usage to memory request (above 1 means the request is too small)
max by (namespace, pod) (
container_memory_working_set_bytes{container="", image="", pod!=""}
)
/
max by (namespace, pod) (
kube_pod_resource_request{resource="memory", unit="bytes"}
)
# Amount over the request (bytes)
max by (namespace, pod) (
container_memory_working_set_bytes{container="", image="", pod!=""}
)
-
max by (namespace, pod) (
kube_pod_resource_request{resource="memory", unit="bytes"}
)
kube_pod_resource_request, so they do not show up in the ratio query. Use unless to find "Pods that have a usage series but no request series". These are the BestEffort Pods, that is, the first candidates to be killed in a global OOM.
# Pods with no memory request, and their actual usage
max by (namespace, pod) (
container_memory_working_set_bytes{container="", image="", pod!=""}
)
unless
max by (namespace, pod) (
kube_pod_resource_request{resource="memory", unit="bytes"}
)
kube_pod_resource_request and kube_node_status_allocatable both have a node label, so these two can be joined without label_replace.
# Per node: sum of memory requests / Allocatable
sum by (node) (
max by (node, namespace, pod) (
kube_pod_resource_request{resource="memory", unit="bytes"}
)
)
/
max by (node) (
kube_node_status_allocatable{resource="memory"}
)
node label, so align labels as in the PromQL above. Using the working set of /kubepods.slice gives the usage of all Pods on the node without summing per Pod.
# Actual usage of all Pods on the node / Allocatable
max by (node) (
container_memory_working_set_bytes{id=~"/kubepods|/kubepods.slice"}
* on(instance) group_left(node) kubelet_node_name
)
/
max by (node) (
kube_node_status_allocatable{resource="memory"}
)
# Maximum memory usage per Pod over the past 7 days (a guide for the request)
max_over_time(
max by (namespace, pod) (
container_memory_working_set_bytes{container="", image="", pod!=""}
)[7d:5m]
)
# 95th percentile over the past 7 days
quantile_over_time(0.95,
max by (namespace, pod) (
container_memory_working_set_bytes{container="", image="", pod!=""}
)[7d:5m]
)
container_memory_rss and decide which to base it on.
Do not base request design on kubectl top
kubectl top pod is convenient, but it is not suited to designing requests. The reason lies in the design of Metrics Server itself.
- What Metrics Server gets from the kubelet is the value of the
/metrics/resourceendpoint, and for memory it is the working set ("Memory is reported as the working set at the instant the metric was collected"). It is the same value ascontainer_memory_working_set_bytesseen through Prometheus, so it includes page cache in the same way - Metrics Server keeps only the latest value. The default collection interval is 60 seconds (
--metric-resolution) and no history is kept, so you cannot look into past peaks - Metrics Server's own README states in its use cases "Don't use Metrics Server when you need: ... An accurate source of resource usage metrics" (README.md)
Use kubectl top to get a feel for "what is using a lot right now", and when you decide a request, specify a period and check with PromQL for comparing requests and actual usage.
Adjust requests automatically with VPA
Instead of running the PromQL in the previous section by hand, you can have the Vertical Pod Autoscaler (VPA) compute recommendations. The idea is the same (derive from quantiles of past usage), and VPA runs it continuously, combined with OOM detection.
VPA consists of three components. Start by looking at recommendations with You would not want Pods to be rebuilt abruptly, so it is safe to first have it only compute recommendations with Check the recommendation as follows. Switch to automatic application Once you confirm the recommendation is reasonable, change This is a configuration example that automatically adjusts only the memory request. With Behavior when an OOM occurs For memory, VPA has a mechanism that raises the recommendation when it detects an OOM kill. The defaults are: From VPA 1.7 onward, these can be specified per container in The memory recommendation itself is computed from quantiles of a usage histogram. The Recommender defaults are: So by default, the 90th percentile of 8 days of peak values plus a 15 % margin becomes the recommended request. The idea is the same as the Warning: The usage VPA looks at is also the working set. The Recommender gets usage from the Metrics Server API ( So for containers with a lot of logging or file I/O, VPA's recommendation comes out larger than the real need. For containers where the gap between Note: Operational points when using VPA:VPA configuration examples and defaults
Component
Role
Recommender
Computes recommendations from usage history and writes them to
status.recommendation of the VPA object
Updater
Evicts Pods that have drifted from the recommendation and prompts their recreation
Admission Controller
Rewrites
resources to the recommendation with a mutating webhook at Pod creationupdateMode: OffupdateMode: Off. In this mode VPA does not touch Pods at all and only writes recommendations to the VPA object. The API definition also says "This can be used for a \"dry run\"".
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: sample-vpa
namespace: default
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: sample
updatePolicy:
updateMode: "Off" # only compute recommendations. Pods are not changed
resourcePolicy:
containerPolicies:
- containerName: app
controlledResources: ["memory"] # default is ["cpu", "memory"]
controlledValues: RequestsOnly # default is RequestsAndLimits
minAllowed:
memory: 128Mi
maxAllowed:
memory: 4Gi
$ kubectl describe vpa sample-vpa
...
Status:
Recommendation:
Container Recommendations:
Container Name: app
Lower Bound:
Memory: 262144k
Target:
Memory: 367001600
Uncapped Target:
Memory: 367001600
Upper Bound:
Memory: 524288k
Target is the recommended request. Lower Bound / Upper Bound are guides for "below this is not enough" and "above this is excessive" respectively, and the Updater makes Pods outside this range eviction targets.updateMode to switch to automatic application. The modes are:
updateModeBehavior
OffOnly computes recommendations. Pods are not changed
InitialApplies recommendations only at Pod creation. Running Pods are not changed
RecreateApplies at creation, and also updates running Pods by evicting and recreating them
InPlaceOrRecreateTries In-Place Resize first, and falls back to recreation if that is not possible. Requires the cluster's
InPlacePodVerticalScaling feature gate
InPlaceTries only In-Place Resize and does not evict. On failure it leaves it to the kubelet's retry. Requires VPA's own
InPlace feature gate in addition to the above
Auto
Deprecated. Currently equivalent to
Recreate
Auto has been deprecated as of VPA 1.7, and the API comment also says "Use explicit update modes like \"Recreate\", \"Initial\", or \"InPlaceOrRecreate\" instead". Do not use Auto in newly written manifests.
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: sample-vpa
namespace: default
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: sample
updatePolicy:
updateMode: "Recreate"
minReplicas: 2 # do not evict when replicas is below this (the global default is also 2)
resourcePolicy:
containerPolicies:
- containerName: app
controlledResources: ["memory"]
controlledValues: RequestsOnly
minAllowed:
memory: 128Mi
maxAllowed:
memory: 4Gi
controlledValues: RequestsOnly, you can manage the limit yourself and leave only the request to VPA. In the context of this article (reducing global OOM and eviction) this is the easy setting to work with. With the default RequestsAndLimits, both are rewritten while keeping the ratio of request to limit, so note that the limit moves too.
Setting
Default
Meaning
oomBumpUpRatio1.2When an OOM kill is detected, record usage at that time multiplied by 1.2 as a sample
oomMinBumpUp
100Mi (104857600)If the bump is smaller than this, bump up by at least this much
memoryAggregationIntervalSeconds
86400 (24 hours)Record one peak sample per this interval
memoryAggregationIntervalCount8How many such intervals to keep. By default the memory recommendation is computed from 24 hours × 8 = 8 days of history
evictAfterOOMSeconds(unset)
Make Pods that OOM within this many seconds of starting eviction targets
containerPolicies.
containerPolicies:
- containerName: app
controlledResources: ["memory"]
oomBumpUpRatio: "1.5" # bump up more strongly on OOM
oomMinBumpUp: 256Mi
memoryAggregationIntervalSeconds: 3600 # record a peak every hour
memoryAggregationIntervalCount: 24 # compute from 24 intervals = 1 day of history
Flag
Default
Use
--target-memory-percentile0.9Quantile used to compute
Target (the recommended request)
--recommendation-lower-bound-memory-percentile0.5Quantile used to compute
Lower Bound
--recommendation-upper-bound-memory-percentile0.95Quantile used to compute
Upper Bound
--recommendation-margin-fraction0.15Safety margin added to the computed value (15 %)
quantile_over_time(0.95, ...) PromQL introduced in the previous section, and it is easiest to think of VPA as running that continuously, combined with OOM detection.metrics.k8s.io/v1beta1), so as described earlier it derives recommendations from values that include page cache.container_memory_rss and container_memory_working_set_bytes is large, it may waste less in the end to cap with maxAllowed or to decide by hand instead of leaving it to VPA.metrics.k8s.io went GA (v1) in v1.37 through KEP-5207, but on the Metrics Server side it is still unsupported even in v0.9.0, the latest release at the time of writing (September 2026), and is being addressed in issue #1786 and PR #1855. That is why this article writes v1beta1. That the referenced value is the working set does not change with the API version.
controlledResources: ["memory"] as in this article's examples, CPU can be left to HPARecreate rebuilds Pods: the Updater evicts Pods using the Eviction API, so set a PodDisruptionBudget. It does not evict when replicas is below minReplicas (the global --min-replicas, default 2, if not specified in VPA)InPlacePodVerticalScaling feature gate
FAQ
My working set is below the limit, but the container was OOMKilled
What memcg OOM compares against the limit is memory.current (container_memory_usage_bytes), not the working set. Furthermore, memory.current reaching the limit does not cause an OOM by itself; the kill happens when reclaim fails to bring usage back under the limit. Also, even if your container is under its limit, it can be killed in a node-wide memory exhaustion (global OOM). First use The two OOM kill paths to tell the path.
If none of the values seems to have reached the limit, suspect a spike shorter than the scrape interval. A spike in memory usage shorter than Prometheus' scrape interval (the global scrape_interval defaults to 1 minute; see Configuration) does not show up in metrics. Batch processing right after startup, or suddenly accepting a very large request, are examples.
In that case, check values like these:
| Metric | Emitted by | Description |
|---|---|---|
kube_pod_container_status_last_terminated_reason{reason="OOMKilled"} |
kube-state-metrics | The last termination reason. The value does not count up when OOMs repeat, so from the second time on, detect it by combining with the increase of kube_pod_container_status_restarts_total (also kube-state-metrics) |
container_oom_events_total |
kubelet | Number of OOM kills. But when the container terminates the cgroup and its series disappear, so it is limited to checking cases where only a child process was OOM killed while the container continued |
node_vmstat_oom_kill |
node-exporter | Number of OOM kills on the node. It cannot identify the Pod or container, but is not affected by container recreation |
Besides metrics, also check Last State in kubectl describe pod (exit code 137 means termination by SIGKILL).
Node usage computed from MemAvailable is low, but Pods are evicted
MemAvailable counts part of the page cache as "free", but the kubelet is working-set based and treats all of active_file as "in use". With workloads that do a lot of file I/O or logging, active_file can grow to several GiB, so even if the MemAvailable-based utilization looks low, the kubelet side may be approaching its threshold.
Conversely, on nodes where page cache has not built up much, node-exporter's side comes out higher. Which is higher depends on the environment, so you cannot judge from just one of them. For the breakdown of factors and the PromQL to get the kubelet-side utilization, see Why node-exporter and kubelet usage disagree.
kubectl top pod and the Prometheus working set differ
kubectl top pod shows the sum of the containers' working sets obtained from the kubelet (via Metrics Server). Prometheus' container_memory_working_set_bytes comes from the same source, but it is as stale as the scrape interval, so instantaneous values differ.
Memory working set keeps growing
First check whether container_memory_rss or container_memory_cache is the one growing.
-
container_memory_rssis growing: suspect a memory leak in the application -
container_memory_cacheis growing: page cache is accumulating. But it includes not only the cache of files on disk but also tmpfs and shared memory. To tell which is growing, checkshmeminmemory.staton the node. tmpfs (such asemptyDirwithmedium: Memory) is not reclaimed on nodes without swap, so accumulation leads to OOM kills.shmemsits on the anon LRU, so it does not appear ininactive_fileand is not subtracted from the working set. That means the increase pushes the working set up as is
For containers with no memory limit, page cache accumulates with the whole node's memory as the ceiling. Set a memory limit on containers that handle large files.
The RSS my application reports differs from the RSS metric
They measure different things. The value called "RSS" in Go, Java, Node.js and so on is per-process RSS, whereas container_memory_rss is anonymous pages charged to the container's cgroup.
| Target | Source | Backing value | What it includes |
|---|---|---|---|
process.memoryUsage().rss in Node.js, the RSS column of ps and top
|
Language runtime or tools such as ps (per process) |
VmRSS in /proc/[pid]/status
|
RssAnon + RssFile + RssShmem
|
container_memory_rss |
cAdvisor (embedded in kubelet; per cgroup) |
anon in cgroup v2 memory.stat
|
Anonymous pages only |
There are two main differences:
-
Whether file mappings are included: a process's RSS includes the pages of the executable, shared libraries and
mmaped files that are in physical memory (RssFile).container_memory_rssdoes not. So comparing the same single process, the RSS the application reports is larger byRssFileandRssShmem -
Unit of aggregation: a process's RSS is for one process.
container_memory_rssis per container cgroup, so it is the total of all processes in the container. Also, a process's RSS counts shared pages in each process, whereas in a cgroup a page is charged to only one cgroup
Note that JVM Runtime.totalMemory() and JMX heap usage are not RSS to begin with; they are the heap size. Metaspace, code cache, thread stacks and GC bookkeeping areas are not included, so you cannot compare these values directly with the memory limit.
Neither value is used in the OOM kill decision. What memcg OOM compares against the limit is memory.current (container_memory_usage_bytes). Global OOM occurs when the memory of the whole node is exhausted and does not use any individual container's usage as a threshold.
Which metric should I use for node memory utilization
Use them according to purpose.
| Purpose | Metric to look at | Emitted by |
|---|---|---|
| Get a rough sense of the node's headroom |
node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes
|
node-exporter |
| Judge whether eviction will occur |
container_memory_working_set_bytes{id="/"} / kube_node_status_capacity{resource="memory"}
|
cAdvisor (embedded in kubelet) / kube-state-metrics |
| Judge whether a Pod can be placed | The sum of kube_pod_resource_request{resource="memory"} (the sum of requests, not usage) |
kube-scheduler /metrics/resources
|
Note: What emits
kube_pod_resource_requestis kube-scheduler, not kube-state-metrics (resources.go). kube-state-metrics also has a similarly namedkube_pod_container_resource_requests, but kube-state-metrics itself says in its help text that it recommends using the more accurate kube-scheduler metrics. Unless Prometheus is configured to scrape kube-scheduler's/metrics/resources,kube_pod_resource_requestis not collected.
Upcoming changes
These are in-progress Kubernetes changes that relate to the behavior described in this article.
KEP-2371: cAdvisor-less, CRI-full Container and Pod Stats
KEP-2371 moves the source of container and node metrics from the kubelet's embedded cAdvisor to the container runtime (CRI). Its goal is to remove the situation where the kubelet collects duplicate statistics from both cAdvisor and CRI.
- The
/metrics/cadvisorendpoint and thecontainer_*metric names are kept, and the plan is to change only where the values come from to CRI - The feature gate
PodAndContainerStatsFromCRIwas Alpha in v1.23 and becomes Beta in v1.37, but even at Beta it is disabled by default
So the definition of the working set and the eviction decision method described in this article will not change for the time being. If this migration becomes the default in the future, the source of values becomes the runtime side, so small differences may arise from differences in how cgroups are read.
KEP-2570: Support Memory QoS with cgroups v2
KEP-2570 reflects requests.memory into cgroup v2 memory protection settings. When enabled, the memory for the request is protected from reclaim, and throttling kicks in before the limit is reached.
The feature gate MemoryQoS has been Alpha (disabled by default) since v1.22 and became Beta and enabled by default in v1.37. However, just enabling the feature gate does not make requests.memory reflected in the cgroup. To reflect it, the kubelet needs the following settings, both disabled by default as of v1.37.
-
memoryReservationPolicy: TieredReservation: setsrequests.memoryintomemory.minfor Guaranteed andmemory.lowfor the rest. A setting added to KubeletConfiguration in v1.36, with a default ofNone. WithNone,memory.minis not set -
memoryThrottlingFactor: setsmemory.hightorequests.memory + factor × (limits.memory - requests.memory). Up to v1.36 the default was0.9, but the feature gate itself was disabled by default then, so it had no actual effect. In v1.37 the default isnil, and unless you set it explicitlymemory.highis not set
Both can be confirmed in the KubeletConfiguration type definition.
// MemoryThrottlingFactor specifies the factor multiplied by the memory limit or node allocatable memory
// ...
// Default: nil
MemoryThrottlingFactor *float64 `json:"memoryThrottlingFactor,omitempty"`
// MemoryReservationPolicy controls how the kubelet applies cgroup v2 memory protection.
// "None" (default): The kubelet does not set memory.min for containers and pods,
// ...
MemoryReservationPolicy MemoryReservationPolicy `json:"memoryReservationPolicy,omitempty"`
So even on a v1.37 cluster, requests.memory is not reflected in the cgroup, as in the table above. "Not reflected" is only about the default configuration, and it is reflected if you set the two above.
What is actually written with the default configuration is as in the implementation. When memoryReservationPolicy is None, an explicit 0 is written instead of requests.memory (equivalent to no protection).
if memoryRequest != 0 && m.memoryReservationPolicy == kubeletconfiginternal.TieredReservationMemoryReservationPolicy {
...
} else {
unified[cm.Cgroup2MemoryMin] = "0"
unified[cm.Cgroup2MemoryLow] = "0"
}
So with the v1.37 default configuration, a container's cgroup has memory.min = 0, memory.low = 0 and memory.high unset (max), and only memory.max (= limits.memory) is in effect. If you use a managed service, check the provider's release notes to see whether these settings get enabled.
What the cgroup v2 values become if you enable it
This section works out what gets written to each cgroup v2 file when you explicitly set memoryThrottlingFactor: 0.9, the default up to v1.36, and also enable memoryReservationPolicy: TieredReservation.
The kubelet configuration looks like this. Suppose you create the following Pod in this state. Assume the node's Allocatable memory is 16Gi. The So this container is throttled with strong reclaim pressure at about 947 MiB, before it reaches 1Gi and gets OOM killed. Up to 256Mi it is also less likely to be reclaimed thanks to Organized by QoS class, it looks like this. What matters is that even with the same For a BestEffort Pod on the same node as the example above (Allocatable 16Gi), You can check how much is protected across the whole node with the following metrics the kubelet exposes (both ALPHA as of v1.37). Warning: If you enable If you are considering enabling it, the precondition is to first right-size requests with Align memory request with actual usage.kubelet configuration example and the values written for each QoS class
MemoryQoS is enabled by default in v1.37, so you can omit featureGates, but I include it to make the intent explicit.
# /var/lib/kubelet/config.yaml (KubeletConfiguration)
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
featureGates:
MemoryQoS: true # true by default in v1.37
memoryReservationPolicy: TieredReservation # default is None
memoryThrottlingFactor: 0.9 # default in v1.37 is nil (= memory.high is not set)
apiVersion: v1
kind: Pod
metadata:
name: sample
spec:
containers:
- name: app
image: nginx
resources:
requests:
memory: 256Mi
limits:
memory: 1Gi
requests.memory and limits.memory differ, so this Pod's QoS class is Burstable. The following values are written to the container's cgroup.
File
Value
Calculation
memory.max
1073741824 (1Gi)
limits.memory as is (set as before, independent of MemoryQoS)
memory.low
268435456 (256Mi)Burstable, so
requests.memory goes to memory.low
memory.min0
requests.memory goes in only for Guaranteed. Otherwise an explicit 0 is written
memory.high
993210368 (about 947 MiB)floor((256Mi + (1Gi - 256Mi) × 0.9) / 4096) × 4096memory.high calculation is rounded down to a multiple of the page size (4 KiB), as in the implementation.
memoryHigh = int64(math.Floor(
float64(memoryRequest)+
(float64(memoryLimitVal)-float64(memoryRequest))*float64(*m.memoryThrottlingFactor))/float64(defaultPageSize)) * defaultPageSize
memory.low.requests / limits, the files that get set change with the QoS class.
QoS class
memory.minmemory.lowmemory.high
Guaranteed (
requests = limits)requests.memory0
Not set (stays
max; excluded when requests and limits are equal)
Burstable (
requests < limits)0requests.memoryfloor((req + (lim − req) × factor) / page size) × page size
Burstable (no
limits)0requests.memorySame formula using the node's Allocatable instead of
limits
BestEffort (no
requests/limits)00floor(Allocatable × factor / page size) × page sizememory.high is floor(16Gi × 0.9 / 4096) × 4096 = 15461879808 (about 14.4 GiB). memory.high is also set on containers with no limit, so enabling memoryThrottlingFactor changes how page cache accumulates as well.memory.min and memory.low are set not only on the container cgroup but also on higher levels (qos_container_manager_linux.go). The kernel evaluates protection by walking up through ancestor cgroups, so the same protection is needed above as well.
Level
memory.minmemory.low
/kubepods.sliceSum of Guaranteed requests + sum of Burstable requests
Sum of Burstable requests
/kubepods.slice/kubepods-burstable.slice0Sum of Burstable requests
Pod cgroup
The Pod's request if Guaranteed
The Pod's request if Burstable
Container cgroup
As in the table above
As in the table above
Metric
Emitted by
Description
kubelet_memory_qos_node_memory_min_byteskubelet
Total reserved as
memory.min for Guaranteed Pods. The amount of memory the kernel will never reclaim
kubelet_memory_qos_node_memory_low_byteskubelet
Total reserved as
memory.low for Burstable PodsmemoryReservationPolicy: TieredReservation, the sum of Guaranteed Pods' requests is hard-reserved on the node as memory.min. The kernel cannot reclaim the memory.min portion, so on nodes with many Guaranteed Pods with oversized requests, the reclaimable memory shrinks and the node may become less stable. The reason the default of memoryReservationPolicy is None is also written in the type definition comment: "This is the default to maintain node stability by preventing \"locked\" memory."
Hands-on verification with kind
These two sections reproduce the behavior described above on a local kind cluster.
Verifying it with kind
You can verify everything so far on your machine by bringing up a v1.37 cluster locally with kind. The following configuration file creates an environment with MemoryQoS enabled. If you omit You read cgroup files by going into the node container. kind puts the cgroup root at Warning: Two behaviors cannot be reproduced in kind. One is the root cgroup fallback described earlier. A kind node is a container and has its own cgroup namespace, so the The other is eviction. kind's kubelet overrides These are measured values on a node with Allocatable of 2005512Ki (= 2053644288 bytes) with Each Warning: Setting only memory to Also, besides the real container, a Pod has a pause (sandbox) container's Protection is also set on higher levels, not only on the container cgroup. On The kubelet metrics described earlier can also be obtained from this endpoint. Steps to bring up a local v1.37 cluster with kind and verify
# kind-memqos.yaml
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
name: memqos
nodes:
- role: control-plane
# the kind default image changes with the version, so pin it explicitly
image: kindest/node:v1.37.0@sha256:a1ed56cfb0e7b93589bdf97c8cd566405a265939e3620fc4f5de89adff580ae5
kubeadmConfigPatches:
- |
kind: KubeletConfiguration
featureGates:
MemoryQoS: true
memoryReservationPolicy: TieredReservation
memoryThrottlingFactor: 0.9
image, the default image embedded in kind is used, but that changes with each kind version. The digest above is the same as the default of kind v0.33.0; to try a different Kubernetes version, replace it with the digest listed in the kind release notes.
$ kind create cluster --config kind-memqos.yaml
# Check that the settings reached the kubelet
$ kubectl get --raw "/api/v1/nodes/memqos-control-plane/proxy/configz" | jq '.kubeletconfig | {memoryThrottlingFactor, memoryReservationPolicy, featureGates}'
{
"memoryThrottlingFactor": 0.9,
"memoryReservationPolicy": "TieredReservation",
"featureGates": { "MemoryQoS": true }
}
/kubelet.slice/kubelet-kubepods.slice, so note that the path differs from /kubepods.slice in a real cluster./sys/fs/cgroup visible from inside the node is actually a delegated non-root cgroup. As a result memory.current exists, and the anon + file substitution does not happen. The difference from a real machine is as follows.
# Inside a kind node (root of the cgroup namespace = actually a non-root cgroup)
$ docker exec memqos-control-plane ls /sys/fs/cgroup/memory.current
/sys/fs/cgroup/memory.current
# On the real host (the root cgroup of cgroup v2)
$ colima ssh -- ls /sys/fs/cgroup/memory.current
ls: /sys/fs/cgroup/memory.current: No such file or directory
$ colima ssh -- ls /sys/fs/cgroup/memory.stat
/sys/fs/cgroup/memory.stat
evictionHard with only imagefs.available / nodefs.available / nodefs.inodesFree, so there is no memory.available threshold. As a result memory-triggered eviction does not occur, and Capacity and Allocatable are the same value. To try eviction, set evictionHard explicitly.
# Identify the cgroup by the real container ID (a glob can pick up the pause side, as explained below)
$ CID=$(kubectl get pod sample-burstable \
-o jsonpath='{.status.containerStatuses[0].containerID}' | sed 's|containerd://||')
$ docker exec memqos-control-plane sh -c "
d=\$(find /sys/fs/cgroup/kubelet.slice -type d -name 'cri-containerd-${CID}.scope')
grep . \$d/memory.max \$d/memory.min \$d/memory.low \$d/memory.high"
/sys/fs/cgroup/.../cri-containerd-50fc89e1....scope/memory.max:1073741824
/sys/fs/cgroup/.../cri-containerd-50fc89e1....scope/memory.min:0
/sys/fs/cgroup/.../cri-containerd-50fc89e1....scope/memory.low:268435456
/sys/fs/cgroup/.../cri-containerd-50fc89e1....scope/memory.high:993210368
memoryThrottlingFactor: 0.9 set.
Pod
QoS
memory.minmemory.lowmemory.high
requests: 256Mi / limits: 1Gi
Burstable
0
268435456
993210368
requests: 128Mi / no limitsBurstable
0
134217728
1861701632
memory only
requests = limits = 512MiBurstable
0
536870912
max (not set)
cpu and memory
requests = limits
Guaranteed
536870912
0
max (not set)
nothing specified
BestEffort
0
0
1848279040
memory.high follows the formula floor((req + (lim − req) × factor) / page size) × page size described earlier. For Pods with no limits, the node's Allocatable is used for lim, and for BestEffort req is 0.requests = limits does not make a Pod Guaranteed. Guaranteed requires both cpu and memory to have requests = limits in every container. A Pod with only memory matched is Burstable, so requests.memory goes to memory.low, not memory.min.memory.high stays unset (max) not because of the QoS class but on the condition requests.memory == limits.memory (the memoryRequest != memoryLimitSpec check in the implementation), so even a Burstable Pod with only memory matched stays at max.
memory.min memory.low memory.high
Burstable with only memory matched 0 536870912 max
Guaranteed with cpu and memory matched 536870912 0 max
cri-containerd-*.scope alongside it, so a glob for cri-containerd-*.scope matches two entries. Which comes first depends on the container ID order, so using the first glob result may read the pause side, and memory.low and memory.high can look like 0 / max. Identify it with .status.containerStatuses[].containerID as in the command above.kubelet-kubepods.slice (equivalent to /kubepods.slice in a real cluster), memory.min is the sum of Guaranteed and Burstable requests and memory.low is the sum of Burstable requests.
kubelet-kubepods.slice memory.min 1243611136 memory.low 706740224
kubelet-kubepods-burstable.slice memory.min 0 memory.low 706740224
kubelet-kubepods-besteffort.slice memory.min 0 memory.low 0
1243611136 - 706740224 = 536870912 = 512Mi ← matches the Guaranteed Pod's request
memory_min_bytes covers only the Guaranteed part, so note that its value differs from the root cgroup's memory.min (Guaranteed + Burstable).
$ kubectl get --raw "/api/v1/nodes/memqos-control-plane/proxy/metrics" | grep kubelet_memory_qos
kubelet_memory_qos_node_memory_low_bytes 7.06740224e+08
kubelet_memory_qos_node_memory_min_bytes 5.36870912e+08
Verifying "reclaim comes first even at the limit" with kind
The claim from earlier that " Looking at the container's cgroup, usage is pinned just short of the limit, but no OOM occurred even once. At this point the working set is On the other hand, doing the same with anonymous pages causes an immediate OOM kill. The kernel log has a line containing Note that after an OOM kill the container's cgroup disappears, so the Confirming that no OOM occurs when page cache pins usage at the limit
memory.current reaching memory.max does not cause an OOM" is easy to reproduce with kind. From a container with limits.memory: 256Mi, write 6 times the limit of data to an on-disk emptyDir.
$ kubectl exec memtest -- sh -c 'dd if=/dev/zero of=/disk/big bs=1M count=1536'
1610612736 bytes (1.5GB) copied, 1.501464 seconds, 1023.0MB/s
memory.max 268435456 (256Mi)
memory.high 248299520 ← floor((64Mi + (256Mi - 64Mi) * 0.9) / 4096) * 4096
memory.current 246484992
memory.peak 249049088
--- memory.events ---
low 0
high 2666 ← number of times it was throttled for exceeding memory.high
max 0 ← memory.max was not reached
oom 0
oom_kill 0 ← no OOM kill occurred
oom_group_kill 0
--- memory.stat (excerpt) ---
file 235982848
inactive_file 235937792 ← nearly all of it is reclaimable inactive page cache
active_file 45056
246484992 - 235937792 = 10547200 (about 10 MiB). The relationship is observed directly: even when usage is pinned at the limit, if the contents are reclaimable page cache, the working set stays small and no OOM occurs.tail /dev/zero keeps allocating memory until it finishes reading its input, and on a node with no swap it cannot be reclaimed.
$ kubectl run oomtest --image=busybox:1.37 --restart=Never \
--overrides='{"spec":{"containers":[{"name":"app","image":"busybox:1.37",
"command":["sh","-c","tail /dev/zero"],
"resources":{"limits":{"memory":"64Mi"}}}]}}'
$ kubectl get pod oomtest -o jsonpath='{.status.containerStatuses[0].state.terminated}'
{"exitCode":137,"reason":"OOMKilled",...}
oom_memcg=. As described earlier, cAdvisor's parser matches this format, so the cgroup can be extracted, and as a result no SystemOOM event is recorded.
oom-kill:constraint=CONSTRAINT_MEMCG,nodemask=(null),cpuset=cri-containerd-9be8306a....scope,
mems_allowed=0,oom_memcg=/docker/.../kubelet-kubepods-burstable-pod....slice,
task_memcg=/docker/.../cri-containerd-9be8306a....scope,task=tail,pid=4867,uid=0
Memory cgroup out of memory: Killed process 4867 (tail) total-vm:68880kB,
anon-rss:64768kB, file-rss:1280kB, shmem-rss:0kB, UID:0 pgtables:176kB oom_score_adj:984
$ kubectl get events -A --field-selector reason=SystemOOM
No resources found
container_oom_events_total series itself goes away. It is the same even if you do not let it be recreated with restartPolicy: Never. To trace the number of OOMs afterward, use node_vmstat_oom_kill (per node) or kube_pod_container_status_last_terminated_reason.
Closing thoughts
The takeaway of this article is that the value you casually call "memory usage" is actually several different ones, namely the working set, RSS, memory.current and MemAvailable, and that the value that triggers an OOM kill is different from the value that triggers an eviction.
I often get questions like "the container's memory usage keeps growing, so isn't it leaking memory?" and "it is OOMKilled but I don't know why", and I kept giving similar explanations each time. Eviction and QoS classes are in the official documentation, and the cgroup v2 specification is in the kernel documentation, but I could not find anything that explains in one place what is actually used as a container's memory usage, and why OOM kills and evictions happen as a result, so I decided to write it myself.
Three points are useful to keep in mind when investigating:
- What you see with
kubectl toporcontainer_memory_working_set_bytesis the working set, and it includes active page cache - What memcg OOM compares against the limit is
memory.current(=container_memory_usage_bytes), not the working set. And it does not occur just by reaching the limit; the kill happens when reclaim fails to bring usage back under it - What triggers eviction is free memory computed from the node's working set, not
MemAvailable
References
Components and versions
This article assumes cgroup v2. With cgroup v1, the mapping between metrics and memory.stat fields is different.
The statements in this article target the following versions.
| Component | Version | Repository | Role in this article |
|---|---|---|---|
| Linux kernel | v6.12 | https://github.com/torvalds/linux | cgroup v2 memory control and the OOM killer itself; the OOM message format in kernel logs |
| Kubernetes (kubelet) | v1.37.0 | https://github.com/kubernetes/kubernetes | Eviction decisions, oom_score_adj, the SystemOOM event, counting for container_oom_events_total
|
| Kubernetes (kube-scheduler) | v1.37.0 | https://github.com/kubernetes/kubernetes | Exposing kube_pod_resource_request / kube_pod_resource_limit (/metrics/resources) |
| cAdvisor | v0.57.0 | https://github.com/google/cadvisor | Embedded in the kubelet; exposes container_* metrics; computes the working set |
| containerd | v2.3.5 | https://github.com/containerd/containerd | Decides that a container's termination reason is OOMKilled
|
| node-exporter | v1.12.1 | https://github.com/prometheus/node_exporter | Exposes node_memory_* (/proc/meminfo) and node_vmstat_oom_kill (/proc/vmstat) |
| kube-state-metrics | v2.20.0 | https://github.com/kubernetes/kube-state-metrics | Exposes kube_node_status_capacity and kube_pod_container_status_*
|
| Metrics Server | v0.8.1 | https://github.com/kubernetes-sigs/metrics-server | Source of the values kubectl top reads |
| Vertical Pod Autoscaler | 1.7.1 | https://github.com/kubernetes/autoscaler | Computing and automatically applying request/limit recommendations |
| kind | v0.33.0 | https://github.com/kubernetes-sigs/kind | Used to verify v1.37 behavior locally |
Note: I assume containerd as the container runtime. With other runtimes such as CRI-O the meaning of cgroup v2 values is the same, but the
OOMKilleddetection logic and cgroup path naming differ.I also assume the cluster has Prometheus, node-exporter and kube-state-metrics installed and that you can run PromQL. These are not part of Kubernetes itself, so if you do not have them, substitute
kubectl topor readmemory.statdirectly on the node.
Links
Listed in the order they appear in the article.
-
Memory Resource Controller (cgroup v2) (definitions of
memory.current/memory.max/memory.statand other files) - About cgroup v2 (cgroup v2 from Kubernetes' point of view)
- Resource Management for Pods and Containers (basics of requests and limits)
- Reserve Compute Resources for System Daemons (calculation of Allocatable)
- Node-pressure Eviction (overview of eviction: signals, thresholds, selection of target Pods)
- Concepts overview - OOM killer (conditions and role of the OOM killer invocation)
- Node out of memory behavior (behavior when eviction is not in time)
- Pod Quality of Service Classes (how QoS classes are determined)
- proc.rst - oom_score / oom_score_adj (the score used to select the process to kill)
-
sysctl/vm.rst (OOM-related sysctls such as
panic_on_oom) - Understand Pressure Stall Information (PSI) Metrics (detecting waits for memory allocation)
-
Metrics Server FAQ (definition and limits of the values
kubectl topshows) - Vertical Pod Autoscaler (recommending requests and limits from usage)
Metrics covered in this article
Grouped by emitting component. The takeaway of this article is that even metrics representing the same "memory usage" have different definitions if they come from different emitters, so when investigating, first check which component emits the value.
cAdvisor (kubelet /metrics/cadvisor)
cAdvisor embedded in the kubelet reads cgroups and exposes them. You do not need to run a separate cAdvisor.
| Metric | Backing value in cgroup v2 | Summary |
|---|---|---|
container_memory_usage_bytes |
memory.current |
Usage including everything. What memcg OOM is evaluated against |
container_memory_working_set_bytes |
memory.current - inactive_file |
The value Kubernetes treats as usage. The basis for kubectl top and eviction decisions |
container_memory_rss |
anon in memory.stat
|
Amount of anonymous pages. Its scope differs from true RSS |
container_memory_cache |
file in memory.stat
|
Page cache size (includes shmem) |
container_memory_total_active_file_bytes |
active_file in memory.stat
|
Active page cache. Included in the working set |
container_memory_total_inactive_file_bytes |
inactive_file in memory.stat
|
Inactive page cache. Subtracted from the working set |
container_memory_mapped_file |
file_mapped in memory.stat
|
Amount of mmaped files |
container_spec_memory_limit_bytes |
memory.max |
The container's memory limit. Exposed as 0 when memory.max is max (no limit) |
container_pressure_memory_waiting_seconds_total |
total of some in memory.pressure
|
Cumulative time (seconds) some processes were waiting for memory |
container_pressure_memory_stalled_seconds_total |
total of full in memory.pressure
|
Cumulative time (seconds) all processes were stalled waiting for memory |
The whole node can be selected with id="/", all Pods on the node with id=~"/kubepods|/kubepods.slice", and a single Pod with container="", image="", pod!="". Only id="/" has no memory.current, so its usage is substituted with anon + file from memory.stat (kernel memory is not included). For details see Values the kubelet uses for eviction decisions.
| Metric | Emitted by | Summary |
|---|---|---|
container_oom_events_total |
kubelet (the name says container_ but it is not cAdvisor) |
Number of OOM kills parsed from oom-kill: lines in the kernel log. Exposed on the same /metrics/cadvisor
|
kubelet (/metrics/resource)
An endpoint that exposes only CPU, memory and swap usage per container, per Pod and per node. It is a different path from /metrics/cadvisor, so to scrape it with Prometheus you need to add configuration.
Metrics Server gets values from here, and the official documentation also says "The metrics-server fetches resource metrics from the kubelets" (Resource metrics pipeline). It is not exclusive to Metrics Server, though, and Metrics Server's own README directs you to scrape this endpoint directly for monitoring purposes ("In such cases please collect metrics from Kubelet /metrics/resource endpoint directly.").
| Metric | Summary |
|---|---|
container_memory_working_set_bytes |
Per-container working set (the same value as the metric of the same name on /metrics/cadvisor) |
pod_memory_working_set_bytes |
Per-Pod working set. Obtained from the Pod's cgroup |
node_memory_working_set_bytes |
Per-node working set |
kubelet (/metrics)
| Metric | Summary |
|---|---|
kubelet_node_name |
A gauge whose value is always 1 (the info-metric pattern). The node name is in the node label, so you can bring node into the container_* metrics, which have no node label, with on(instance) group_left(node)
|
kubelet_memory_qos_node_memory_min_bytes |
Total reserved as memory.min for Guaranteed Pods (when MemoryQoS is enabled. ALPHA) |
kubelet_memory_qos_node_memory_low_bytes |
Total reserved as memory.low for Burstable Pods (when MemoryQoS is enabled. ALPHA) |
node-exporter
Turns the contents of /proc directly into metrics. It knows nothing of Kubernetes concepts.
| Metric | Source | Summary |
|---|---|---|
node_memory_MemTotal_bytes |
/proc/meminfo |
Total physical memory on the node |
node_memory_MemAvailable_bytes |
/proc/meminfo |
The kernel's estimate of allocatable memory. Counts reclaimable page cache as free |
node_memory_Cached_bytes |
/proc/meminfo |
Page cache size (includes tmpfs / shmem) |
node_memory_Shmem_bytes |
/proc/meminfo |
The tmpfs / shared memory part of the above |
node_vmstat_oom_kill |
/proc/vmstat |
Number of OOM kills on the node. Includes both memcg OOM and global OOM |
kube-state-metrics
Converts Kubernetes API objects (spec / status) into metrics. It does not handle actual resource usage.
| Metric | Summary |
|---|---|
kube_node_status_capacity{resource="memory"} |
The node's Capacity. The denominator for evaluating eviction thresholds |
kube_node_status_allocatable{resource="memory"} |
The amount assignable to Pods. The upper bound the scheduler uses |
kube_pod_container_status_last_terminated_reason |
The last termination reason (reason="OOMKilled" detects an OOM kill). No series appears for a container that has never terminated
|
kube_pod_container_status_restarts_total |
Container restart count |
kube_pod_container_resource_requests |
The request written in the manifest (per container) |
kube_pod_container_resource_limits |
The limit written in the manifest (per container) |
kube-scheduler (/metrics/resources)
Exposed on the secure port (default 10259) at /metrics/resources. It is a different endpoint from /metrics.
| Metric | Summary |
|---|---|
kube_pod_resource_request{resource="memory", unit="bytes"} |
The sum of requests per Pod. No series is emitted for a Pod whose value is 0 |
kube_pod_resource_limit{resource="memory", unit="bytes"} |
The sum of limits per Pod |
Top comments (0)