In production Kubernetes clusters running microservices under high concurrency, few operational failures are more frustrating than sudden Exit Code 137 (OOMKilled) cascades and relentless CrashLoopBackOff cycles.
Standard observability metrics often show container memory hovering at 70% before the Linux kernel abruptly executes a SIGKILL. In this architectural deep-dive, we deconstruct how the kernel cgroup v2 memory controller enforces limits, why anonymous vs active file page reclaim causes unexpected kills, and how to systematically profile native memory leaks.
1. Demystifying Exit Code 137 & cgroup v2 Enforcement
When Kubernetes terminates a pod with Exit Code 137, it indicates the process received signal 9 (128 + 9 = 137). This is enforced at the Linux kernel level via control groups:
[Container Memory Usage]
│
▼
[memory.current] ──► [memory.high Throttle Window] ──► [memory.max Hard Ceiling]
│ (Kernel OOM Invocation)
▼
[SIGKILL (Exit Code 137)]
▼
[K8s CrashLoopBackOff]
Under cgroup v2:
-
memory.current: Actual total memory consumed (Anonymous memory + Page cache + Kernel memory). -
memory.high: The throttling threshold where the kernel slows down allocation threads before killing. -
memory.max: The hard ceiling. If exceeded and page cache reclaim fails, the kernel's OOM killer fires immediately.
2. The Three Common Culprits Behind Unexpected OOMKilled
A. JVM & Node.js Native / Off-Heap Memory Leaks
Setting -Xmx4g on a 5Gi memory limit container is insufficient. Native off-heap allocations (DirectByteBuffers, thread stacks, JNI, glibc malloc fragmentation) bypass JVM heap counters:
# Inspect native memory tracking in running JVM pod
$ jcmd 1 VM.native_memory baseline
$ jcmd 1 VM.native_memory detail.diff
B. Linux Kernel Page Cache Reclaim Stall
When containers process gigabytes of streaming I/O or log files, page cache expands rapidly. If the pages are marked dirty or the disk I/O subsystem suffers backpressure, the kernel cannot reclaim buffer pages fast enough, mistaking active cache for anonymous bloat.
C. Go Runtime pprof Heap vs OS RSS Mismatch
In Go microservices, runtime.MemStats.HeapAlloc might show 500MB, while ps aux shows RSS exceeding 2GB. This is caused by MADV_DONTNEED vs MADV_FREE scavenging delays on older kernels.
3. Production Hardening: Quality of Service (QoS) & Manifest Tuning
To prevent noisy neighbors from stealing buffer headroom, configure Guaranteed QoS where requests equal limits:
apiVersion: apps/v1
kind: Deployment
metadata:
name: payment-settlement-engine
spec:
replicas: 4
template:
spec:
containers:
- name: settlement-worker
image: internal-registry/settlement:v2.4
resources:
requests:
cpu: "2000m"
memory: "4Gi"
limits:
cpu: "2000m"
memory: "4Gi"
env:
- name: JAVA_TOOL_OPTIONS
value: "-XX:MaxRAMPercentage=75.0 -XX:+UseG1GC -XX:+ExitOnOutOfMemoryError"
4. Diagnostics & Kernel cgroup Inspection Commands
Run these inside the container or host node to diagnose the memory footprint:
# 1. Inspect container cgroup v2 memory breakdown
cat /sys/fs/cgroup/memory.stat
# 2. Check OOM kill event count in cgroup v2
cat /sys/fs/cgroup/memory.events
# 3. Stream kernel dmesg for invocation logs
dmesg -T | grep -i -E "oom-killer|killed process"
🛠️ Free Zero-Trust Engineering Utilities
If you are generating Kubernetes manifests or analyzing network and system configs, check out our free browser-based developer tools:
- Docker Run to Kubernetes YAML Converter: https://global-utils.com/en/docker-to-k8s
- Zero-Trust HAR Sanitizer & Visualizer: https://global-utils.com/en/har-analyzer
- Systems Engineering Playbooks & Guides: https://global-utils.com/en/blog
Originally published at NerdKit Engineering.
Top comments (0)