A Kubernetes resource request reserves scheduling capacity. A low CPU reading beside a large request can be worth investigating, but it does not establish how much a cloud invoice will fall if you change that request.
I maintain kube-saver, an MIT-licensed Python CLI for exploring that gap. It collects Kubernetes resource requests and available metrics-server samples, models CPU and memory costs using configurable rates, and writes local HTML reports and resource-change plans. It also has a terminal dashboard.
This walkthrough uses v2.0.0 and focuses on interpreting a report before making a change.
Install the reviewed version
You need Python 3.10 or newer and a reachable Kubernetes cluster with the documented read permissions.
python3 -m venv .venv
source .venv/bin/activate
python -m pip install "kube-saver==2.0.0"
kube-saver version
Version 2.0.0 is published on PyPI. The pinned command above selects the version used in this walkthrough; kube-saver version should print kube-saver 2.0.0.
A wheel is also available from the v2.0.0 GitHub release. If you use that route, download the wheel and SHA256SUMS.txt, compare the wheel's SHA256 with its checksum entry, and install the downloaded wheel in your virtual environment.
kube-saver does not require a hosted kube-saver account or cloud billing account. Cluster authentication is still necessary. A kubeconfig credential plugin may contact its identity provider.
Check the connection, then collect a report
Replace staging-cluster below with the actual kubeconfig context you intend to inspect:
export KUBE_SAVER_CONTEXT=staging-cluster
kube-saver doctor --context staging-cluster
kube-saver report -o cost-report.html --json cost-report.json
doctor checks context, connectivity, permissions, and Metrics API availability. Missing metrics can produce optional warnings without preventing a request-based report. A passing diagnostic does not guarantee fresh measurements for every pod.
The report command contacts the Kubernetes API. Once generated, the self-contained HTML file can be viewed offline. JSON supports further local processing. Reports include workload and namespace names, so inspect them before sharing. For interactive exploration, run kube-saver without a subcommand.
Read the cost figures as a model
The pricing implementation uses a 730-hour month and separate CPU and memory rates:
monthly modeled cost =
730 × (
CPU millicores / 1000 × USD per CPU-core-hour
+ memory bytes / 1024³ × USD per GiB-hour
)
This prices the resource quantity being considered: requested allocation, a measured request–usage gap, or a proposed request reduction. Bundled provider rates are assumptions, not a live rate feed. Configure rates appropriate to your environment.
An illustrative calculation, not a benchmark or measurement: at chosen rates of $0.04 per CPU-core-hour and $0.005 per GiB-hour, pricing 500m CPU and 1 GiB memory gives:
730 × (0.5 × 0.04 + 1 × 0.005) = $18.25 per modeled month
The model does not ingest cloud invoices. Shared-node packing, discounts, autoscaler behavior, storage, networking, and other charges can make realized billing changes different from a CPU/memory estimate. Lower requests can create room for another workload without reducing billed node count.
Missing metrics do not prove a pod is idle
metrics-server supplies current CPU and memory samples when available and valid. Invalid, incomplete, stale, or future-dated samples are treated as unavailable; the default maximum sample age is 300 seconds.
Without valid usage data, kube-saver falls back to estimates and suppresses actionable right-sizing recommendations. A large displayed waste figure in estimate mode is not evidence of measured idleness. Read source and degraded-coverage indicators before drawing conclusions. A partial scan cannot establish the behavior of pods it did not collect.
Treat recommendations as candidates
The engine uses snapshots, not historical peaks or workload-specific percentile sizing. CLI defaults apply:
- CPU headroom of 1.5 times observed usage and memory headroom of 1.2 times observed usage.
- Absolute minimums of 100m CPU and 128Mi memory.
- Normal-mode floors of half the current CPU and memory requests.
- Upward rounding to whole millicores and MiB.
Absolute floors remain in aggressive mode; aggressive mode skips relative floors.
Controller consolidation considers all collected siblings, including busy replicas that would not independently generate a candidate, and takes the largest suggestion across those samples. Estimated usage, excluded siblings, and multi-container siblings suppress the workload plan. Differing sibling requests suppress that resource's recommendation because the active template during a rollout cannot safely be inferred from those observations.
These checks reduce particular mistakes; they do not prove a workload will tolerate a change. Review burst behavior, startup requirements, memory peaks, SLOs, exclusions, and rollout state. StatefulSets and PVC-backed workloads are not automatically protected. Confidence labels describe utilization ratios rather than statistical guarantees.
Generate a local review plan
kube-saver pr-plan -d ./pr-files
This writes summary.md, review.txt, README.md, and apply-patches.sh. It does not open a GitHub PR or apply Kubernetes changes.
For GitOps, translate reviewed request changes into the manifests your reconciler manages. Use your normal branch, code review, staging validation, and rollout monitoring. Keep reports containing private cluster names out of public commits.
The generated script provides an alternative explicit apply path. Executing it contacts Kubernetes and may trigger a rollout. It requires KUBE_SAVER_APPLY_CONTEXT, passes that context to each patch, and stops on the first failure. Earlier successful patches are not rolled back. Review every patch and its target context. Avoid direct patches that your GitOps reconciler will immediately overwrite.
A useful first report gives you a scoped question: which workloads have a measured request–usage gap, and which have insufficient evidence to size safely? Gather representative-load history before reducing requests, then compare operational behavior after a reviewed change.
The documentation covers installation, safety, RBAC, and configuration. The repository contains the implementation and contribution guide.
Top comments (0)