DEV Community

Cover image for kube-memtier: Measuring Per-Pod Cold Memory with DAMON
Raihan
Raihan

Posted on

kube-memtier: Measuring Per-Pod Cold Memory with DAMON

Introduction

How much of your Kubernetes pods' memory is actually cold? RAM is expensive, and Kubernetes only knows how much memory a pod holds, not how much it uses. Cold memory could live on cheaper tiers (compressed RAM via zswap, NVMe swap, CXL memory) without hurting performance, but nobody can see it per pod today. With Kubernetes swap going GA in v1.34, the question "which pods could safely swap, and how much?" now has a real, practical answer.

kube-memtier: Measuring Per-Pod Cold Memory

kube-memtier is a Linux/Kubernetes node agent that measures, per pod, how much memory is hot (in use right now), warm (idle for a little while) and cold (untouched for minutes). It uses the Linux kernel's own access monitor, DAMON, and exports the results as Prometheus metrics.

The Problem

Cold memory is invisible today. Kubernetes knows how much memory a pod requests and holds, but has no visibility into how much is cold and could safely move to cheaper tiers. With Kubernetes swap going GA in v1.34, the question "which pods could safely swap, and how much?" now has a real, practical answer.

The Solution

kube-memtier measures per-pod memory temperature using the kernel's DAMON and exports Prometheus metrics:

Metric What It Measures
memtier_lru_hot_bytes active_anon + active_file (kernel LRU baseline)
memtier_lru_cold_bytes inactive_anon + inactive_file
memtier_damon_hot_bytes bytes accessed in the last DAMON aggregation window
memtier_damon_warm_bytes idle bytes, idle for less than --cold-after
memtier_damon_cold_bytes bytes idle for at least --cold-after

Getting Started

Prerequisites

  • Linux kernel ≥ 6.8 (newer is better; DAMON gains features every release)
  • cgroup v2 mounted at /sys/fs/cgroup
  • Go ≥ 1.22 (for building the agent)

Installation

# 1. Install Go (≥ 1.22)
sudo apt update && sudo apt install -y golang-go

# 2. Clone the repo
git clone https://github.com/raihan-js/kube-memtier.git
cd kube-memtier

# 3. Build and verify
make check          # gofmt + go vet + unit tests (MUST pass)
make build workload # builds bin/memtier-agent and bin/hotcold

# 4. First experiment (needs sudo)
sudo bash scripts/run-hotcold-experiment.sh 1024 128 5m vaddr first-run

# 5. Analyze results
python3 scripts/analyze.py experiments/results/$(ls -d experiments/results/ | head -1)
Enter fullscreen mode Exit fullscreen mode

Key Features

  • LRU mode: No root needed, coarse baseline from kernel's memory.stat
  • DAMON mode: Per-cgroup memory monitoring via kernel's DAMON sysfs
  • Prometheus metrics: Full metric family exposition at /metrics
  • Grafana dashboard: Ready-to-use 6-panel dashboard for monitoring
  • Experiment framework: CSV-driven sweep script for RQ1/RQ2 research questions

Experiment Protocol

The project includes a comprehensive experiment protocol (~400 runs, 25-30 machine-hours):

Variable Values
total MiB 1024, 4096
hot fraction 5%, 12.5%, 25%, 50%
ops vaddr, paddr+memcg
sample / aggr 5ms/50ms, 5ms/100ms, 10ms/200ms
nr_regions max 100, 1000
warmup 2 × cold-after
reps 5

Hypotheses tested:

  • H1: Without memory pressure, LRU baseline underestimates cold anon memory substantially
  • H2: DAMON vaddr overestimates total bytes (holes) but hot estimate is close to truth
  • H3: Error grows when hot fraction is small and max_nr_regions is low
  • H4: kdamond CPU stays < 1% of a core for max_nr_regions ≤ 1000 at 5ms sampling

Architecture

The project is organized into well-defined packages with clear boundaries:

Package Responsibility Must NOT
internal/cgroup Read-only cgroup v2: memory.stat parsing, pod cgroup discovery, PIDs Write anything
internal/damon DAMON sysfs: start kdamond, snapshot tried_regions, stop Contain policy (what counts as cold)
internal/estimator Pure functions: classify regions → Breakdown{Hot,Warm,Cold,IdleAtLeast}. No I/O. Do I/O
internal/exporter Prometheus text format rendering, /metrics handler Know about memory
cmd/memtier-agent Flags, modes, loop, JSON for --once Contain parsing logic

Gotchas and Caveats

Please read the risks documentation before running experiments:

  • vaddr may not exist on your kernel - Some kernels list only paddr in avail_operations
  • paddr monitored nothing without an explicit region - The agent sets the biggest "System RAM" range from /proc/iomem (root needed)
  • sysfs layout drifts between kernels - filters/ vs core_filters/ vs ops_filters/ varies by version
  • Only one DAMON user at a time - If damo or another tool owns a kdamond, writes fail with EBUSY
  • vaddr region bytes ≠ resident bytes - Regions can span unmapped holes between mappings
  • Age is an approximation of idle time, and starts at 0 - Warmup ≥ 2 × cold-after
  • Sampling resolution and merging - DAMON checks one page per region per sample
  • Transparent huge pages - One touched byte marks a whole 2 MiB page young

Contributing

We welcome contributions! Please see the CODE OF CONDUCT and CONTRIBUTING guidelines.

How to Contribute

  1. Fix bugs or add features - open an issue first
  2. Add experiment configurations - expand the CSV sweep matrix
  3. Add new metrics - following the existing metric naming conventions
  4. Improve documentation - update AGENTS.md, CHANGELOG.md, or the experiment protocol
  5. Run the experiment matrix and share your results

Development Setup

# Clone the repo
git clone https://github.com/raihan-js/kube-memtier.git
cd kube-memtier

# Install dependencies
go get ./...

# Run tests
make check

# Build the agent
make build workload

# Run the first experiment (needs sudo)
sudo bash scripts/run-hotcold-experiment.sh 1024 128 5m vaddr first-run

# Analyze results
python3 scripts/analyze.py experiments/results/$(ls -d experiments/results/ | head -1)
Enter fullscreen mode Exit fullscreen mode

Cloud Native & Linux Foundation

This project is well-suited for CNCF (Cloud Native Computing Foundation) and Linux Foundation involvement:

  • CNCF relevance: Measures memory per pod, a key concern for Kubernetes cluster management
  • Linux kernel involvement: DAMON is a Linux kernel feature; contributions can go to kernel mailing lists
  • Kubernetes integration: Fits naturally with Kubernetes swap and memory management
  • Measurement infrastructure: The experiment framework and analysis scripts are reusable

To contribute to CNCF/LF:

  1. File a CNCF project proposal
  2. Submit patches to the Linux kernel DAMON maintainers
  3. Join the kube-memtier community
  4. Present at KubeCon + CloudNative Con

License

Apache-2.0 (the CNCF default). Contributions need a DCO sign-off (git commit -s).

References


Originally developed on Zorin OS (Ubuntu-based) with Ryzen 5 5600G, 32 GB RAM, RTX 3060. Tested on Linux 6.18 with DAMON vaddr mode.

This project is part of the kube-memtier initiative to measure per-pod memory temperature using kernel DAMON.

Top comments (0)