Originally published on the Professional IT Services blog.
Twelve queries show what a Kubernetes cluster costs and where it wastes money: nodes and taints, requests and limits, real usage, top pods, QoS classes, idle requests, pods over their requests, volumes, orphans, leftovers, load balancers and restarts. They need only kubectl, jq and the hcloud CLI, and the output below comes from our production K3s cluster.
That is the most direct route to reducing Kubernetes operational overhead: before changing anything, measure what is requested, what is used and what is billed. Our earlier post covers what we changed. This one is the checklist we run first, and the same queries run on client clusters before their cost reviews.
All results were measured read-only on 1 October 2026 on pits-prod-fsn1-cluster-01: K3s on Hetzner Cloud, five nodes, 98 running pods, 135 containers. Client namespaces are anonymised (client-a, client-b, shop-wp), pod-hash suffixes are trimmed, and long outputs are cut to the rows that matter.
The 12 queries at a glance
| # | Query | Question it answers | Our result |
|---|---|---|---|
| 1 | kubectl get nodes -o custom-columns |
What are we paying for, and what is tainted? | 5 nodes, 42 vCPU, 4 of 5 with an empty instance-type label |
| 2 | kubectl describe nodes |
How much is requested and limited per node? | CPU limits 124% of capacity, 204% on one node |
| 3 | kubectl top nodes |
What is actually used? | 4.7% CPU, one node at 14% |
| 4 | kubectl top pods -A --sort-by=memory |
Who uses the memory? | One database pod at 4,290 Mi |
| 5 | QoS census with jq
|
Who is evicted first? | 0 Guaranteed, 77 Burstable, 21 BestEffort |
| 6 | Requests against usage per pod | Where are requests idle? | About 2.7 GiB idle on two database replicas |
| 7 | Pods above their requests | Where is the other side of the problem? | 12 pods above a request |
| 8 | Volume capacity against use | Which volumes are empty? | 14 of 16 10-GB volumes under 1% used |
| 9 | Orphaned volumes | What do we pay for with no pod attached? | None |
| 10 | Zero-replica workloads, leftover jobs | What was left behind? | 1 StatefulSet at 0, 10 completed job pods |
| 11 |
LoadBalancer and NodePort Services |
What is billed outside the cluster? | One Service, one billed load balancer |
| 12 | Restarts and OOMKills | Is anything already breaking? | No OOMKill |
Why run the queries before every review, not just once?
Because the numbers move without anyone deciding to move them. In May 2026 we published how we cut Kubernetes resource overhead by 50% using only built-in tools, which took the CPU limits commitment from 105% down to 77.7%. Five months later the same metric reads 124% across the cluster and 204% on one node.
There was no single bad decision behind it. The drift has two parts:
| May 2026 post | 1 October 2026 | |
|---|---|---|
| CPU limits committed | 77.7% | 124% |
| Cluster capacity | about 58 vCPU | 42 vCPU |
| CPU limits in cores | about 45 | 52.05 |
| CPU limits against the old 58 vCPU | n/a | about 90% |
- The denominator shrank. After Hetzner raised its prices in spring 2026 (usage in April and May is already billed at the new rates: CAX11 from €3.29 to €4.49, CAX31 from €11.99 to €15.99, CAX41 from €23.99 to €31.49), we went looking for ways to downscale. The worker count stayed at four, but capacity went from 58 to 42 vCPU. Monthly compute cost went from €114.95 to €81.95, which is €33 or 29% less. The newest node, a CX53 (16 vCPU, 32 GB) at €29.49 a month, replaced a CAX41 at €31.49 on 29 September.
- The limits grew. Based on usage metrics, we raised the limit-to-request ratio on many deployments, the "low requests, generous limits" pattern from the May post.
On the old 58 vCPU the same 52.05 cores of limits would read about 90%, so most of the jump is the smaller denominator and the rest is deliberate. The May capacity is inferred from the Hetzner invoices, not measured, so treat 58 and 90% as approximate. The lesson stands either way: run the queries before every review, because a good number in May says nothing about October.
The queries, with real output
1. Node inventory: type, architecture, capacity, taints
kubectl get nodes -o custom-columns='NAME:.metadata.name,TYPE:.metadata.labels.node\.kubernetes\.io/instance-type,ARCH:.status.nodeInfo.architecture,CPU:.status.allocatable.cpu,TAINTS:.spec.taints[*].key'
hcloud server list -o columns=name,type,location
NAME TYPE ARCH CPU TAINTS
master-node k3s arm64 2 <none>
worker-node-6 <none> amd64 8 dedicated
worker-node-7 <none> amd64 8 <none>
worker-node-8 <none> amd64 8 <none>
worker-node-9 <none> amd64 16 <none>
NAME TYPE LOCATION
master-node cax11 fsn1
worker-node-6 cx43 fsn1
worker-node-8 cx43 fsn1
worker-node-7 cx43 fsn1
worker-node-9 cx53 fsn1
What it tells you: the inventory, and the first gap: K3s without the Hetzner cloud controller manager leaves the instance-type label empty (k3s on the master, <none> elsewhere), so the price mapping needs hcloud server list. Read the TAINTS column now; it explains query 3.
2. Requests and limits per node
kubectl describe nodes | awk '/^Name:/{n=$2} /^ (cpu|memory) /{print n, $1, $5}'
master-node cpu (95%)
master-node memory (26%)
worker-node-6 cpu (48%)
worker-node-6 memory (45%)
worker-node-7 cpu (138%)
worker-node-7 memory (118%)
worker-node-8 cpu (204%)
worker-node-8 memory (117%)
worker-node-9 cpu (117%)
worker-node-9 memory (113%)
The awk keeps only the limits percentage from each "Allocated resources" block. For the cluster total against allocatable CPU:
alloc=$(kubectl get nodes -o json | jq '[.items[].status.allocatable.cpu | tonumber] | add * 1000')
kubectl get pods -A --field-selector=status.phase=Running -o json | jq -r --argjson a "$alloc" '
def cpu: if . == null then 0 elif endswith("m") then (.[:-1] | tonumber) else (tonumber * 1000) end;
def sum(f): [.items[].spec.containers[].resources | f | cpu] | add;
sum(.requests.cpu) as $r | sum(.limits.cpu) as $l
| "requests \($r)m (\($r * 100 / $a | round)%), limits \($l)m (\($l * 100 / $a | round)%)"'
requests 8425m (20%), limits 52050m (124%)
What it tells you: how far the node can be oversold. Memory limits above 100% mean that if every pod grew to its limit, the node would run out of memory and the kernel would OOM-kill pods or the kubelet would evict them. CPU above 100% means throttling, not a crash.
3. Actual usage
kubectl top nodes --no-headers | awk '{c+=$2; m+=$4} END {printf "%dm CPU, %.1f GiB\n", c, m/1024}'
kubectl top node worker-node-6 --no-headers | awk '{print $1, "cpu", $2, "mem", $5}'
1968m CPU, 25.3 GiB
worker-node-6 cpu 45m mem 14%
| CPU | Memory | |
|---|---|---|
| Allocatable | 42,000m | 80.1 GiB |
| Requested | 8,425m (20%) | 34.2 GiB (43%) |
| Actually used | 1,968m (4.7%) | 25.3 GiB (32%) |
| Used as a share of requested | 23% | 74% |
| Limits committed | 52,050m (124%) | 78.7 GiB (98%) |
What it tells you: the gap between what was reserved and what runs. CPU is the wide one: 20% requested, 4.7% used. The outlier is worker-node-6: 45m CPU and 2.2 GiB of memory, 14% of an 8 vCPU / 15.2 GiB node that costs €15.99 a month.
Query 1 explains why. MariaDB Galera is no longer pinned to one dedicated node: its replicas spread over three labelled workers, one per node. But worker-node-6 still carries the dedicated=mariadb:NoSchedule taint, so nothing except Galera and the DaemonSets can land there. Today it runs one Galera replica plus four DaemonSet pods. The chart comment says the taint stays "until a multi-arch audit", and since 29 September every worker is amd64. The taint is a leftover from the mixed-architecture period, and it keeps every other workload off the node. Re-check that the reason behind each taint still holds.
4. Top pods by memory
kubectl top pods -A --sort-by=memory --no-headers | head -6 | awk '{print $1"/"$2, $4}'
prod-mariadb/mariadb-galera-new-0 4290Mi
prod-mariadb/mariadb-galera-new-1 1441Mi
prod-mariadb/mariadb-galera-new-2 1272Mi
prod-prometheus/prometheus-prometheus-stack-kube-prom-prometheus-0 1255Mi
prod-mailu/mailu-clamav-0 985Mi
prod-prometheus/loki-chunks-cache-0 881Mi
What it tells you: where the memory goes. The database primary dominates, and the Loki chunks cache is down to 881 Mi, so the change from the May post held (it was 8 GB). Sort by cpu the same way; Prometheus used 243m.
5. QoS census
kubectl get pods -A --field-selector=status.phase=Running -o json | jq -r '.items[].status.qosClass' | sort | uniq -c
kubectl get pods -A --field-selector=status.phase=Running -o json | jq -r '[.items[].spec.containers[]] as $c | "\($c | map(select(.resources.limits.memory == null)) | length) of \($c | length) containers have no memory limit"'
21 BestEffort
77 Burstable
56 of 135 containers have no memory limit
What it tells you: who is evicted first under memory pressure. No Guaranteed line means none (that class needs requests equal to limits on every container). The 21 BestEffort pods include all of cert-manager, the hcloud CSI controller and node plugins, and the k3s svclb pods; 56 of 135 containers have no memory limit at all.
6. Requests against usage, per pod
Both this query and the next one start from one joined table: requests from the pod specs, usage from the metrics API.
kubectl get pods -A --field-selector=status.phase=Running -o json > pods.json
kubectl get --raw /apis/metrics.k8s.io/v1beta1/pods > usage.json
jq -rs '
def q: if . == null then 0 else
capture("^(?<n>[0-9.]+)(?<u>[A-Za-z]*)$") as $c
| ($c.n | tonumber) * ({"n":1e-9,"u":1e-6,"m":1e-3,"":1,"Ki":1024,"Mi":1048576,"Gi":1073741824}[$c.u]) end;
(.[1].items | map({(.metadata.namespace + "/" + .metadata.name):
{cpu: ([.containers[].usage.cpu | q] | add), mem: ([.containers[].usage.memory | q] | add)}}) | add) as $u
| .[0].items[] | (.metadata.namespace + "/" + .metadata.name) as $k | select($u[$k])
| [$k, ([.spec.containers[].resources.requests.cpu | q] | add) * 1000, $u[$k].cpu * 1000,
([.spec.containers[].resources.requests.memory | q] | add) / 1048576, $u[$k].mem / 1048576]
| map(if type == "number" then floor else . end) | @tsv' pods.json usage.json > usage.tsv
# columns: pod, cpu request (m), cpu usage (m), memory request (Mi), memory usage (Mi)
awk -F'\t' '{print $1 "\t" $4 "\t" $5 "\t" $4-$5}' usage.tsv | sort -t$'\t' -k4 -nr | head -4
prod-mariadb/mariadb-galera-new-2 4096 1272 2824
prod-mariadb/mariadb-galera-new-1 4096 1441 2655
prod-nextcloud/nextcloud-… 2176 652 1524
prod-mailu/mailu-clamav-0 2048 985 1063
What it tells you: the biggest idle memory requests (request, usage, difference in Mi): two database replicas hold roughly 2.6 and 2.8 GiB they do not use. Swap columns 4 and 5 for 2 and 3 to rank CPU: Nextcloud requests 300m and uses 13m, shop-wp requests 300m and uses 28m, and six Mailu and ingress pods request 200m each while using 1 to 8m. A database buffer pool is meant to be warm, so read this as a list of where to look, not what to cut.
7. Pods that use more than they request
awk -F'\t' '($4>0 && $5>$4) || ($2>0 && $3>$2) {n++} END {print n " pods above a request"}' usage.tsv
awk -F'\t' '$4>0 && $5>$4 {print $1 "\t" $4 "\t" $5 "\t" $5-$4}' usage.tsv | sort -t$'\t' -k4 -nr | head -3
12 pods above a request
prod-mariadb/mariadb-galera-new-0 4096 4290 194
prod-prometheus/prometheus-prometheus-stack-kube-prom-prometheus-0 1074 1257 183
ingress-nginx/ingress-nginx-controller-… 128 268 140
What it tells you: the other side of over-requesting. Under-requested pods are the first evicted when a node runs short of memory. Besides these memory rows, Dovecot used 327m CPU against a 200m request, and all three ingress-nginx pods were above their memory request.
8. Volume capacity against actual use
for n in $(kubectl get nodes -o name | cut -d/ -f2); do
kubectl get --raw /api/v1/nodes/$n/proxy/stats/summary | jq -r '.pods[].volume[]? | select(.pvcRef) | [.pvcRef.namespace + "/" + .pvcRef.name, .capacityBytes, .usedBytes] | @tsv'
done | sort -u > pvc.tsv
awk -F'\t' '$2 > 9e9 && $2 < 11e9 {n++; if ($3 / $2 < 0.01) u++} END {print n " volumes of ~10 GB, " u " under 1% used"}' pvc.tsv
hcloud volume list -o noheader -o columns=size | awk '{s+=$1} END {print NR " volumes, " s " GB"}'
16 volumes of ~10 GB, 14 under 1% used
24 volumes, 510 GB
What it tells you: the kubelet stats endpoint (it needs permission on nodes/proxy) reports capacity and use per claim. 14 of the 16 ten-gigabyte volumes were under 1% used, most under 60 MB, yet 10 GB is Hetzner's minimum volume size (Hetzner's documentation gives the range as 10 GB to 10 TB), so the saving is fewer volumes, not smaller ones. 510 GB at €0.0572 per GB-month is about €29 a month; the July invoice shows €31.22 for volumes, about the price of two CX43 nodes. One gotcha: the two local emptydir-storage claims report the node's root disk (149.9 GiB), not the claim.
9. Orphaned volumes
kubectl get pv -o json | jq -r '.items[] | select(.status.phase != "Bound") | .metadata.name'
comm -23 \
<(kubectl get pvc -A -o json | jq -r '.items[] | .metadata.namespace + "/" + .metadata.name' | sort) \
<(kubectl get pods -A -o json | jq -r '.items[] | .metadata.namespace as $ns | .spec.volumes[]? | select(.persistentVolumeClaim) | $ns + "/" + .persistentVolumeClaim.claimName' | sort -u)
(no output)
What it tells you: the first command lists Released and Available volumes, the second lists claims that no pod mounts. Both were empty. A clean result is worth keeping in the review: the Hetzner CSI volumes use the Delete reclaim policy, so a deleted claim takes its billed volume with it.
10. Zero-replica workloads and leftover jobs
kubectl get deploy,sts -A -o json | jq -r '.items[] | select(.spec.replicas == 0) | .kind + " " + .metadata.namespace + "/" + .metadata.name'
kubectl get pods -A --field-selector=status.phase=Succeeded --sort-by=.metadata.creationTimestamp \
-o custom-columns='NAMESPACE:.metadata.namespace,POD:.metadata.name,CREATED:.metadata.creationTimestamp'
StatefulSet prod-nextcloud/nextcloud-redis-replicas
NAMESPACE POD CREATED
prod-mailu backup-29723220-… 2026-07-07T03:00:00Z
... (9 more completed job pods)
What it tells you: what was paused or finished and never removed. One StatefulSet sat at zero replicas, and there were 10 completed job pods, the oldest a July mail-backup run. A zero-replica workload is a deliberate pause or a forgotten one, and the volumes it owns would still be billed, which is what query 9 catches.
11. LoadBalancer and NodePort Services
kubectl get svc -A -o json | jq -r '.items[] | select(.spec.type == "LoadBalancer" or .spec.type == "NodePort") | [.metadata.namespace, .metadata.name, .spec.type, (.status.loadBalancer.ingress[0].ip // "-"), ([.spec.ports[].nodePort] | join(","))] | @tsv'
hcloud load-balancer list -o columns=name,type
hcloud primary-ip list -o noheader -o columns=type | sort | uniq -c
ingress-nginx ingress-nginx-controller LoadBalancer <master-ip> 30221,31065
NAME TYPE
load-balancer-1 lb11
5 ipv4
5 ipv6
What it tells you: kubectl shows one LoadBalancer Service, answered by the k3s ServiceLB on the master's IP. The invoice shows a Hetzner Load Balancer 11 (€7.49 a month) forwarding ports 80 and 443 to NodePorts 30221 and 31065 on four nodes. It is not an orphan, only invisible from the Service. The five primary IPv4 addresses are the same story: separate invoice lines (€0.50 each) that no Kubernetes object points to.
12. Restarts and OOMKills
kubectl get pods -A -o json | jq -r '.items[] | .metadata.namespace as $ns | .metadata.name as $p | .status.containerStatuses[]? | select(.lastState.terminated.reason == "OOMKilled") | [$ns, $p, .name] | @tsv'
kubectl get pods -A -o json | jq -r '.items[] | .metadata.namespace as $ns | .metadata.name as $p | .status.containerStatuses[]? | select(.restartCount > 0) | [$ns + "/" + $p, .name, .restartCount, (.lastState.terminated.finishedAt // "-")] | @tsv' | sort -t$'\t' -k3 -nr | head -8
(no OOMKilled rows)
kube-system/hcloud-csi-node-… hcloud-csi-driver 14 2026-06-17T19:19:06Z
...
prod-mailu/mailu-front-… front 1 2026-09-29T12:40:59Z
prod-mailu/mailu-front-… front 1 2026-09-29T12:41:59Z
What it tells you: whether a tight limit is already hurting. There was no OOMKill anywhere. The 14 restarts of the CSI node driver date from the 17 June outage, and the two mailu-front restarts from 29 September.
What can't kubectl tell you about the bill?
Everything above is a view of the cluster. The invoice is a view of the cloud account, and three things sit between them:
- Instance types. Without the cloud controller manager the label is empty, so a node cannot be priced from kubectl alone.
- Load balancers. A balancer created outside the cluster has no Service of its own. The only trace in Kubernetes is a NodePort.
- Addresses and volumes. Primary IPv4 addresses are billed per address, and every PVC becomes a billed volume with a 10 GB floor.
So map each node, Service and volume to an invoice line once, by hand, and keep the table with the queries. Ours, from the July 2026 invoice (net prices):
| What kubectl shows | Invoice line |
|---|---|
| 5 nodes | CAX11 €4.49, CAX41 €31.49, 3 × CX43 €15.99 each (the CAX41 has since become a CX53 at €29.49) |
ingress-nginx-controller Service |
Load Balancer 11, €7.49 |
| Node public addresses | 5 × Primary IPv4, €0.50 each |
| 24 PVCs on Hetzner volumes | Volume, €0.0572 per GB-month, 545.761 GB-months, €31.22 |
Compute, load balancer, addresses and volumes came to €125.28 a month on that invoice, and €134.96 with the Storage Box and Object Storage. That is the number the 12 queries are trying to move.
How we use these queries in practice
Every deployment on our clusters starts with sensible requests and limits, and gets tuned from the metrics afterwards. The queries above are the metrics part. We run the same set on client clusters before their cost reviews, because the order of work matters: measure first, then change requests, then look at nodes and volumes.
For a team that wants this done once, that is a cloud cost optimization audit: the same read of requests, limits, usage and invoice lines, delivered as a prioritised action plan. If the numbers should stay healthy (limits drift back, a taint outlives its reason, a node swap changes the denominator), it belongs in a Fractional DevOps retainer. To see your own cluster's numbers first, book a free discovery call.
Top comments (0)