Introduction
How much of your Kubernetes pods' memory is actually cold? RAM is expensive, and Kubernetes only knows how much memory a pod holds, not how much it uses. Cold memory could live on cheaper tiers (compressed RAM via zswap, NVMe swap, CXL memory) without hurting performance, but nobody can see it per pod today. With Kubernetes swap going GA in v1.34, the question "which pods could safely swap, and how much?" now has a real, practical answer.
kube-memtier: Measuring Per-Pod Cold Memory
kube-memtier is a Linux/Kubernetes node agent that measures, per pod, how much memory is hot (in use right now), warm (idle for a little while) and cold (untouched for minutes). It uses the Linux kernel's own access monitor, DAMON, and exports the results as Prometheus metrics.
The Problem
Cold memory is invisible today. Kubernetes knows how much memory a pod requests and holds, but has no visibility into how much is cold and could safely move to cheaper tiers. With Kubernetes swap going GA in v1.34, the question "which pods could safely swap, and how much?" now has a real, practical answer.
The Solution
kube-memtier measures per-pod memory temperature using the kernel's DAMON and exports Prometheus metrics:
| Metric | What It Measures |
|---|---|
memtier_lru_hot_bytes |
active_anon + active_file (kernel LRU baseline) |
memtier_lru_cold_bytes |
inactive_anon + inactive_file |
memtier_damon_hot_bytes |
bytes accessed in the last DAMON aggregation window |
memtier_damon_warm_bytes |
idle bytes, idle for less than --cold-after |
memtier_damon_cold_bytes |
bytes idle for at least --cold-after |
Getting Started
Prerequisites
- Linux kernel ≥ 6.8 (newer is better; DAMON gains features every release)
- cgroup v2 mounted at
/sys/fs/cgroup - Go ≥ 1.22 (for building the agent)
Installation
# 1. Install Go (≥ 1.22)
sudo apt update && sudo apt install -y golang-go
# 2. Clone the repo
git clone https://github.com/raihan-js/kube-memtier.git
cd kube-memtier
# 3. Build and verify
make check # gofmt + go vet + unit tests (MUST pass)
make build workload # builds bin/memtier-agent and bin/hotcold
# 4. First experiment (needs sudo)
sudo bash scripts/run-hotcold-experiment.sh 1024 128 5m vaddr first-run
# 5. Analyze results
python3 scripts/analyze.py experiments/results/$(ls -d experiments/results/ | head -1)
Key Features
- LRU mode: No root needed, coarse baseline from kernel's memory.stat
- DAMON mode: Per-cgroup memory monitoring via kernel's DAMON sysfs
-
Prometheus metrics: Full metric family exposition at
/metrics - Grafana dashboard: Ready-to-use 6-panel dashboard for monitoring
- Experiment framework: CSV-driven sweep script for RQ1/RQ2 research questions
Experiment Protocol
The project includes a comprehensive experiment protocol (~400 runs, 25-30 machine-hours):
| Variable | Values |
|---|---|
| total MiB | 1024, 4096 |
| hot fraction | 5%, 12.5%, 25%, 50% |
| ops | vaddr, paddr+memcg |
| sample / aggr | 5ms/50ms, 5ms/100ms, 10ms/200ms |
| nr_regions max | 100, 1000 |
| warmup | 2 × cold-after |
| reps | 5 |
Hypotheses tested:
- H1: Without memory pressure, LRU baseline underestimates cold anon memory substantially
- H2: DAMON vaddr overestimates total bytes (holes) but hot estimate is close to truth
- H3: Error grows when hot fraction is small and max_nr_regions is low
- H4: kdamond CPU stays < 1% of a core for max_nr_regions ≤ 1000 at 5ms sampling
Architecture
The project is organized into well-defined packages with clear boundaries:
| Package | Responsibility | Must NOT |
|---|---|---|
internal/cgroup |
Read-only cgroup v2: memory.stat parsing, pod cgroup discovery, PIDs | Write anything |
internal/damon |
DAMON sysfs: start kdamond, snapshot tried_regions, stop | Contain policy (what counts as cold) |
internal/estimator |
Pure functions: classify regions → Breakdown{Hot,Warm,Cold,IdleAtLeast}. No I/O. | Do I/O |
internal/exporter |
Prometheus text format rendering, /metrics handler | Know about memory |
cmd/memtier-agent |
Flags, modes, loop, JSON for --once | Contain parsing logic |
Gotchas and Caveats
Please read the risks documentation before running experiments:
-
vaddr may not exist on your kernel - Some kernels list only
paddrinavail_operations -
paddr monitored nothing without an explicit region - The agent sets the biggest "System RAM" range from
/proc/iomem(root needed) -
sysfs layout drifts between kernels -
filters/vscore_filters/vsops_filters/varies by version -
Only one DAMON user at a time - If
damoor another tool owns a kdamond, writes fail withEBUSY - vaddr region bytes ≠ resident bytes - Regions can span unmapped holes between mappings
- Age is an approximation of idle time, and starts at 0 - Warmup ≥ 2 × cold-after
- Sampling resolution and merging - DAMON checks one page per region per sample
- Transparent huge pages - One touched byte marks a whole 2 MiB page young
Contributing
We welcome contributions! Please see the CODE OF CONDUCT and CONTRIBUTING guidelines.
How to Contribute
- Fix bugs or add features - open an issue first
- Add experiment configurations - expand the CSV sweep matrix
- Add new metrics - following the existing metric naming conventions
- Improve documentation - update AGENTS.md, CHANGELOG.md, or the experiment protocol
- Run the experiment matrix and share your results
Development Setup
# Clone the repo
git clone https://github.com/raihan-js/kube-memtier.git
cd kube-memtier
# Install dependencies
go get ./...
# Run tests
make check
# Build the agent
make build workload
# Run the first experiment (needs sudo)
sudo bash scripts/run-hotcold-experiment.sh 1024 128 5m vaddr first-run
# Analyze results
python3 scripts/analyze.py experiments/results/$(ls -d experiments/results/ | head -1)
Cloud Native & Linux Foundation
This project is well-suited for CNCF (Cloud Native Computing Foundation) and Linux Foundation involvement:
- CNCF relevance: Measures memory per pod, a key concern for Kubernetes cluster management
- Linux kernel involvement: DAMON is a Linux kernel feature; contributions can go to kernel mailing lists
- Kubernetes integration: Fits naturally with Kubernetes swap and memory management
- Measurement infrastructure: The experiment framework and analysis scripts are reusable
To contribute to CNCF/LF:
- File a CNCF project proposal
- Submit patches to the Linux kernel DAMON maintainers
- Join the kube-memtier community
- Present at KubeCon + CloudNative Con
License
Apache-2.0 (the CNCF default). Contributions need a DCO sign-off (git commit -s).
References
- DAMON documentation: https://docs.kernel.org/admin-guide/mm/damon/usage.html
- cgroup v2 documentation: https://docs.kernel.org/admin-guide/cgroup-v2.html
- Kubernetes swap GA: https://kubernetes.io/docs/concepts/swap/
- Prometheus metrics exposition format: https://prometheus.io/docs/practices/metrics/
- Experiment protocol: docs/05-research-and-experiments.md
Originally developed on Zorin OS (Ubuntu-based) with Ryzen 5 5600G, 32 GB RAM, RTX 3060. Tested on Linux 6.18 with DAMON vaddr mode.
This project is part of the kube-memtier initiative to measure per-pod memory temperature using kernel DAMON.
Top comments (0)