Kubernetes swap changes what memory pressure does to a pod: a workload that would have been OOMKilled can remain alive and run slowly instead. That sounds like a pure improvement until you look at what decides which pods get that treatment. Under LimitedSwap, the kubelet and container runtime use the container's memory request when determining its swap allowance. The scheduler uses the same field for placement and does not account for swap.
This post covers one configuration and stays inside it: cgroup v2 nodes running LimitedSwap, and the pods that mode allows to use swap. It extends the model in Kubernetes Requests vs Limits rather than replacing it. That model still describes the default node. This is what changes when you turn swap on.

One number, two jobs: the memory request places the pod and sizes its swap allowance.
Kubernetes Swap Is Off by Default, and Opt-In Per Node
On a default Kubernetes node, workloads cannot use swap. The kubelet refuses to start on a node with swap enabled unless you set failSwapOn: false, and even then it tells the container runtime to allocate zero swap to workloads. Letting pods use it is a separate, explicit choice: set memorySwap.swapBehavior to LimitedSwap in the kubelet configuration, as described in the Kubernetes swap documentation. Different nodes can use different behaviors, which makes this a node pool decision, and each node reports its swap capacity in its status, so you can see which pools have swap provisioned. It sits in the platform layer of modern infrastructure architecture and is felt by every team whose pods land there.
The mechanics matter for what follows. The kubelet does not enforce swap itself. It directs the container runtime to apply the setting, and on cgroup v2 that setting is memory.swap.max on the container's cgroup. On cgroup v1, Kubernetes workloads are not allowed to use swap at all.
Scope of this post:
- Covered: cgroup v2 nodes running LimitedSwap, non-high-priority Burstable pods, the request-derived swap allowance, and the shift from OOMKill to latency.
- Not covered: eviction interaction, swap-aware scheduling, benchmark or density numbers, and which release moved the feature to which stage.
- Basis: behavior as documented on kubernetes.io and in KEP-2400. Nothing here was measured.
The Memory Request Becomes the Swap Entitlement
The requests-vs-limits model treats the memory request as a scheduling signal. The scheduler adds it to a node's accounting ledger and places the pod. Under LimitedSwap the same number gets a second job. The amount of swap a container may use is proportional to its memory request, the node's physical memory, and the total swap available to pods on that node. Raise the request and the allowance grows with it. Lower it and the allowance shrinks.
| Consumer | What it does with the memory request |
|---|---|
| kube-scheduler | Places the pod on a node with enough unreserved memory in its accounting. Does not account for swap. |
| kubelet and container runtime | Derive the container's swap allowance from the request and apply it to the container's cgroup. |
Per the Kubernetes documentation, the scheduler does not currently account for swap usage, and the docs name that as a factor that heightens noisy-neighbor risk. A node can be packed on requests alone while its pods compete for the same swap device. Kubernetes swap gives the same request a second runtime consequence: where a pod lands is still one decision, but the request also helps determine how much swap the container may use when memory runs short.
An OOMKill Is Loud. Swap Is Quiet.
On a default node, a container that hits its memory limit is killed by the OOM killer, the pod reports OOMKilled, and the restart counter moves. It is an event, and most monitoring is built around events.

An OOMKill is an event. Swap pressure can leave a pod running with no event attached.
Swap changes the ending. The Kubernetes documentation says swap can help prevent pods from being terminated during memory pressure spikes. The same page says that having swap available reduces predictability, that swapping data back into memory can be slower by many orders of magnitude, and that swap changes a system's behavior under memory pressure. Both statements describe one mechanism. The pod that survives may now be waiting on a much slower storage path.
How slow depends on the storage behind the swap. The docs state that performance on a node with swap depends on the underlying physical storage, and is significantly worse in an IOPS-constrained environment, such as a cloud VM with I/O throttling, than on SSD or NVMe. The pod is the same either way. The outcome depends on a device it has no say in.
Swap also changes who pays. The docs warn that enabling it increases noisy-neighbor risk, because pods that use their RAM frequently can cause other pods to swap. Kubernetes swap can turn a pod-level memory-pressure event into contention for a shared node storage resource.
⚠ A silent failure mode: A pod that would have been OOMKilled may now be alive and slow. If your alerting keys on OOMKilled status and restart counts, swap removes the signal without removing the problem.
The signals move to places most dashboards do not look. The kubelet exposes container_swap_usage_bytes and container_swap_limit_bytes on its resource metrics endpoint, and kubectl top pod --show-swap reports per-pod swap use. None of that appears in pod status the way an OOMKill does.
Who Gets Swap, and How Much
Eligibility for Kubernetes swap is narrower than "pods on a swap-enabled node." Under LimitedSwap, only non-high-priority pods in the Burstable QoS class may use swap, and BestEffort and Guaranteed pods are prohibited, according to the Kubernetes 1.32 swap post. KEP-2400 scopes the feature to Burstable pods.
| Pod class | Swap under LimitedSwap |
|---|---|
| Guaranteed | None |
| BestEffort | None |
| Burstable, high priority | None |
| Burstable, other | Proportional to the memory request |

Eligibility under LimitedSwap is narrower than pods on a swap-enabled node.
The documentation decides what counts as high priority, so check it for the version you run. There is also a built-in opt-out: a container in a Burstable pod whose memory request equals its memory limit gets no swap. Request sizing is now a swap decision as well as a scheduling decision.
That creates a tension the default requests-vs-limits model does not capture. Sizing requests from steady-state usage is sound for placement. Under LimitedSwap it also means a container with a modest request gets a modest allowance, however spiky its peaks. Anything that sets the memory request is now also setting the potential swap allowance, whether that request comes from a chart default or an autoscaler's recommendation covered in VPA vs HPA.
Before You Enable It: One Question and Four Checks
Diagnostic: "For this workload, is a slowdown on a device I am not monitoring better than a restart I would have seen?"
If the answer is yes for a class of workload, swap is a reasonable tool for it. If the answer is no, setting the request equal to the limit opts the container out.
Before enabling LimitedSwap:
- Confirm the node pool runs cgroup v2. On cgroup v1, workloads cannot use swap.
- Decide which workloads may page. The request sets the allowance, high-priority pods are excluded, and request equal to limit opts a container out.
- Move alerting beyond OOMKilled. Watch container swap usage against its limit, and per-pod swap in kubectl top.
- Check what backs the swap. The docs tie performance to the storage underneath, single out IOPS-constrained environments as the worst case, and strongly encourage encrypting the swap space.
Nothing here argues against swap. Kubernetes swap is a reasonable tool when the trade is chosen deliberately. What it argues against is enabling it as a node setting and leaving requests and alerts exactly where they were. The scheduler will keep placing pods using requests, while the kubelet and container runtime apply the swap allowances derived from those requests.
Architect's Verdict
Kubernetes swap does not soften a memory limit. It changes the failure path available to a memory-pressured pod, and the memory request helps determine how much swap the container may use.
The real problem is that two long-standing habits stop lining up. Teams commonly size requests for the scheduler and alert on the OOMKill. Under LimitedSwap the request also helps set a swap allowance that the scheduler does not account for, and the kill can become a slowdown that no restart counter records. Neither change is a bug. Both are documented. The risk is enabling swap as a node setting without revisiting what requests mean and what alerts watch.
An OOMKill announces itself. Swap fails quietly, and only the teams that chose to measure it will notice.
Additional Resources
- Modern Infrastructure & IaC Architecture — the pillar hub for platform-layer decisions like node configuration and runtime behavior
- Kubernetes Requests vs Limits — the scheduler-and-kernel model this post extends, including QoS classes and default OOM behavior
- VPA vs HPA: Why Most Teams Choose the Wrong Autoscaler — how autoscalers change requests, which now also change swap allowances
- Kubernetes Documentation: Swap memory management — official reference for NoSwap, LimitedSwap, metrics and the documented risks
- Kubernetes 1.32: Fresh Swap Features for Linux Users — eligibility rules, the request-proportional allowance and the request-equals-limit opt-out
Originally published at rack2cloud.com
Top comments (0)