September 2026 | ~15 min read
The $170,000 Wake-Up Call
$14,200/month on EKS with 38% average utilization. We cut it to $4,180/month — a 71% reduction — without a single production incident. Here's the full playbook.
We'd been running Kubernetes in production for two years. The cluster worked fine. The applications were stable. But when our finance team pulled the quarterly cloud spend report, the EKS line item was $42,600 for three months — and climbing. Our platform team sat down with kubectl top nodes and the reality hit: we were paying for 12 nodes running 24/7, most of them sitting at 30-40% CPU utilization. Every pod was requesting 4x the resources it actually consumed. Nobody had touched the Cluster Autoscaler configuration since initial setup. And every single node was on-demand pricing.
This wasn't a technology problem. It was an operational one. We had no cost visibility, no right-sizing process, no spot instance strategy, and no incentive for development teams to care about resource efficiency. So we spent six weeks fixing all of it. Here's everything we did, in the order we did it, with the exact configurations and scripts we used.
Table of Contents
- The Numbers That Matter
- The Problem: Death by Over-Provisioning
- Strategy 1: Right-Sizing Pods with VPA
- Strategy 2: Cluster Autoscaler to Karpenter Migration
- Strategy 3: Spot Instances for Non-Critical Workloads
- Strategy 4: Namespace Resource Quotas and Cost Allocation
- Strategy 5: Node Consolidation and Bin-Packing
- Strategy 6: Graviton/ARM Instances
- Monitoring: Cost Visibility That Actually Drives Change
- Results: Before vs After
- ROI Analysis
- Lessons Learned
- Action Items Checklist
- Conclusion
The Numbers That Matter
Before: Over-Provisioned EKS Cluster
EKS Cluster Overview:
- Node count: 12 nodes (m5.2xlarge, on-demand)
- Total cluster capacity: 96 vCPU / 384 GB RAM
- Average CPU utilization: 38%
- Average memory utilization: 42%
- Pod resource requests: 4x actual usage (average)
- Cost allocation: None — single bill, no team visibility
- Autoscaler: Cluster Autoscaler, default config
- Spot instances: 0%
Monthly Cost Breakdown:
- EC2 nodes (12 × m5.2xlarge): $5,904
- EKS control plane: $146
- EBS volumes (12 × 100GB gp3): $288
- NAT Gateway: $450
- ALB / Ingress: $180
- CloudWatch Logs: $520
- Data transfer: $680
- ECR storage + transfer: $120
- Miscellaneous (Route53, S3, etc.): $190
Total Monthly Spend: $14,200 (~$170,400/year)
Effective utilization cost: $5,396 (38% of compute)
Waste: $8,804/month in over-provisioning
After: Optimized EKS Cluster
EKS Cluster Overview:
- Node count: 5 nodes average (mixed types)
- Total cluster capacity: Dynamic — scales with demand
- Average CPU utilization: 72%
- Average memory utilization: 68%
- Pod resource requests: 1.3x actual usage (right-sized)
- Cost allocation: Per-namespace, per-team dashboards
- Autoscaler: Karpenter with consolidation
- Spot instances: 65% of workloads
Monthly Cost Breakdown:
- EC2 nodes (mixed, spot+OD+Graviton): $1,820
- EKS control plane: $146
- EBS volumes (5 × 80GB gp3): $96
- NAT Gateway (optimized routing): $280
- ALB / Ingress: $180
- CloudWatch + Kubecost: $340
- Data transfer (optimized): $420
- ECR storage + transfer: $90
- Miscellaneous: $160
Total Monthly Spend: $4,180 (~$50,160/year)
Monthly Savings: $10,020 (71% reduction)
Annual Savings: $120,240
The Problem: Death by Over-Provisioning
Our EKS cluster had all the classic symptoms of a Kubernetes deployment that nobody had tuned since day one.
Symptom 1: Pod Resource Requests Were Fiction
When we first deployed to Kubernetes, every team followed the same pattern: copy a deployment manifest from the wiki, set requests and limits to "safe" values, and never look at them again. The wiki example used 1 vCPU and 2 GB RAM for requests. Most of our services actually consumed 200-300m CPU and 256-512 MB RAM under peak load.
# What every team was deploying
resources:
requests:
cpu: "1000m" # Requesting 1 full vCPU
memory: "2Gi" # Requesting 2 GB RAM
limits:
cpu: "2000m"
memory: "4Gi"
# What the pods actually used (p95)
# cpu: 280m
# memory: 420Mi
That 1 vCPU request meant the Kubernetes scheduler reserved an entire vCPU for each pod, even though the pod only used 280m on average. With 40+ pods running across the cluster, we were reserving 40+ vCPU but actually consuming about 11. The scheduler thought the cluster was 80% utilized. The nodes were 38% utilized. The gap was pure waste.
Symptom 2: Cluster Autoscaler Was Asleep at the Wheel
Our Cluster Autoscaler was running with default settings. It worked — technically. When all the reserved capacity was exhausted, it would spin up a new node. The problem was the other direction: it took 10+ minutes to decide a node was underutilized, and it wouldn't remove a node if any pod had a restrictive PodDisruptionBudget. Since most of our PDBs required minAvailable: 1, the autoscaler almost never scaled down. Nodes accumulated like barnacles.
Symptom 3: Zero Cost Visibility
When we asked team leads how much their services cost, nobody had any idea. The entire EKS cluster was a single line item on the AWS bill. There were no resource quotas per namespace, no cost allocation tags, and no dashboards showing per-team spend. Without visibility, there was zero incentive to optimize.
Symptom 4: On-Demand Everything
Every single node was an on-demand m5.2xlarge. No spot instances, no Graviton, no right-sized instance types. We were running batch processing jobs, dev/staging environments, and internal tooling on the same premium compute tier as production API servers.
Strategy 1: Right-Sizing Pods with VPA
The first and highest-impact step was right-sizing pod resource requests. This alone freed enough capacity to drop 4 nodes.
Step 1: Audit Current Resource Waste
Before changing anything, we needed hard data. We wrote a script that pulls actual resource usage from the Kubernetes metrics API and compares it to the requests in each pod spec.
#!/usr/bin/env python3
"""
resource_audit.py — Analyze pod resource waste across the cluster.
Compares actual usage (from metrics-server) against requests/limits.
"""
import subprocess
import json
from dataclasses import dataclass
from typing import Optional
@dataclass
class PodResourceReport:
namespace: str
pod_name: str
container: str
cpu_request_m: int
cpu_usage_m: int
cpu_waste_pct: float
mem_request_mb: int
mem_usage_mb: int
mem_waste_pct: float
def run_kubectl(args: list[str]) -> dict:
"""Run a kubectl command and return parsed JSON output."""
result = subprocess.run(
["kubectl"] + args + ["-o", "json"],
capture_output=True, text=True, check=True
)
return json.loads(result.stdout)
def parse_cpu(value: str) -> int:
"""Parse CPU value to millicores."""
if value.endswith("m"):
return int(value[:-1])
elif value.endswith("n"):
return int(value[:-1]) // 1_000_000
else:
return int(float(value) * 1000)
def parse_memory(value: str) -> int:
"""Parse memory value to MB."""
units = {"Ki": 1024, "Mi": 1024**2, "Gi": 1024**3}
for suffix, multiplier in units.items():
if value.endswith(suffix):
return int(int(value[:-len(suffix)]) * multiplier / (1024**2))
return int(int(value) / (1024**2))
def get_pod_metrics() -> dict:
"""Fetch current resource usage from metrics-server."""
metrics = run_kubectl(["top", "pods", "--all-namespaces", "--containers"])
usage_map = {}
for item in metrics.get("items", []):
ns = item["metadata"]["namespace"]
pod = item["metadata"]["name"]
for container in item.get("containers", []):
key = f"{ns}/{pod}/{container['name']}"
usage_map[key] = {
"cpu_m": parse_cpu(container["usage"]["cpu"]),
"mem_mb": parse_memory(container["usage"]["memory"])
}
return usage_map
def audit_resources(exclude_namespaces: Optional[list] = None) -> list[PodResourceReport]:
"""Compare resource requests vs actual usage for all pods."""
if exclude_namespaces is None:
exclude_namespaces = ["kube-system", "kube-node-lease", "kube-public"]
pods = run_kubectl(["get", "pods", "--all-namespaces"])
metrics = get_pod_metrics()
reports = []
for pod in pods.get("items", []):
ns = pod["metadata"]["namespace"]
if ns in exclude_namespaces:
continue
pod_name = pod["metadata"]["name"]
for container in pod["spec"].get("containers", []):
container_name = container["name"]
requests = container.get("resources", {}).get("requests", {})
if not requests:
continue
cpu_req = parse_cpu(requests.get("cpu", "0m"))
mem_req = parse_memory(requests.get("memory", "0Mi"))
key = f"{ns}/{pod_name}/{container_name}"
usage = metrics.get(key, {"cpu_m": 0, "mem_mb": 0})
cpu_waste = ((cpu_req - usage["cpu_m"]) / cpu_req * 100) if cpu_req > 0 else 0
mem_waste = ((mem_req - usage["mem_mb"]) / mem_req * 100) if mem_req > 0 else 0
reports.append(PodResourceReport(
namespace=ns,
pod_name=pod_name,
container=container_name,
cpu_request_m=cpu_req,
cpu_usage_m=usage["cpu_m"],
cpu_waste_pct=round(cpu_waste, 1),
mem_request_mb=mem_req,
mem_usage_mb=usage["mem_mb"],
mem_waste_pct=round(mem_waste, 1),
))
return reports
def print_report(reports: list[PodResourceReport]):
"""Print a formatted waste report sorted by CPU waste."""
reports.sort(key=lambda r: r.cpu_waste_pct, reverse=True)
print(f"\n{'Namespace':<20} {'Pod':<35} {'CPU Req':>8} {'CPU Use':>8} "
f"{'Waste%':>7} {'Mem Req':>8} {'Mem Use':>8} {'Waste%':>7}")
print("─" * 120)
total_cpu_req = 0
total_cpu_use = 0
total_mem_req = 0
total_mem_use = 0
for r in reports:
total_cpu_req += r.cpu_request_m
total_cpu_use += r.cpu_usage_m
total_mem_req += r.mem_request_mb
total_mem_use += r.mem_usage_mb
flag = " *** " if r.cpu_waste_pct > 60 else ""
print(f"{r.namespace:<20} {r.pod_name:<35} {r.cpu_request_m:>6}m {r.cpu_usage_m:>6}m "
f"{r.cpu_waste_pct:>6.1f}% {r.mem_request_mb:>6}MB {r.mem_usage_mb:>6}MB "
f"{r.mem_waste_pct:>6.1f}%{flag}")
print("─" * 120)
cluster_cpu_waste = (total_cpu_req - total_cpu_use) / total_cpu_req * 100
cluster_mem_waste = (total_mem_req - total_mem_use) / total_mem_req * 100
print(f"{'CLUSTER TOTAL':<56} {total_cpu_req:>6}m {total_cpu_use:>6}m "
f"{cluster_cpu_waste:>6.1f}% {total_mem_req:>6}MB {total_mem_use:>6}MB "
f"{cluster_mem_waste:>6.1f}%")
print(f"\nPotential CPU savings: {total_cpu_req - total_cpu_use}m "
f"({cluster_cpu_waste:.0f}% of requested)")
print(f"Potential memory savings: {total_mem_req - total_mem_use}MB "
f"({cluster_mem_waste:.0f}% of requested)")
if __name__ == "__main__":
reports = audit_resources()
print_report(reports)
When we ran this against our cluster, the output was brutal:
Namespace Pod CPU Req CPU Use Waste% Mem Req Mem Use Waste%
────────────────────────────────────────────────────────────────────────────────────────────────────────────────
payments payments-api-7f8b9c6d4-x2k9m 1000m 180m 82.0% 2048MB 310MB 84.9% ***
analytics analytics-worker-5d4f8a3b2-j7n1 1000m 220m 78.0% 2048MB 480MB 76.6% ***
notifications notif-service-6c9e7d5a1-m3p8 1000m 240m 76.0% 2048MB 390MB 81.0% ***
orders orders-api-8a2b4c6d8-k5h2 1000m 310m 69.0% 2048MB 520MB 74.6% ***
users users-api-3e5f7a9b1-n8q4 1000m 290m 71.0% 2048MB 440MB 78.5% ***
inventory inventory-sync-4b6d8e0f2-r1s5 1000m 350m 65.0% 2048MB 680MB 66.8% ***
frontend web-bff-9c1d3e5f7-t4u6 1000m 420m 58.0% 2048MB 720MB 64.8%
gateway api-gateway-2d4f6a8b0-v7w9 1000m 480m 52.0% 2048MB 890MB 56.5%
────────────────────────────────────────────────────────────────────────────────────────────────────────────────
CLUSTER TOTAL 42000m 11280m 73.1% 86016MB 24820MB 71.2%
Potential CPU savings: 30720m (73% of requested)
Potential memory savings: 61196MB (71% of requested)
73% CPU waste. 71% memory waste. Every namespace was over-provisioned by at least 50%.
Step 2: Deploy the Vertical Pod Autoscaler
Rather than manually adjusting every deployment (we had 40+ across 8 namespaces), we deployed VPA to continuously monitor actual usage and recommend right-sized requests.
# vpa/vpa-recommender.yaml
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: payments-api-vpa
namespace: payments
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: payments-api
updatePolicy:
# Start with "Off" to get recommendations without auto-applying
# Switch to "Auto" after validating recommendations for 1-2 weeks
updateMode: "Off"
resourcePolicy:
containerPolicies:
- containerName: payments-api
minAllowed:
cpu: "100m"
memory: "128Mi"
maxAllowed:
cpu: "2000m"
memory: "4Gi"
controlledResources: ["cpu", "memory"]
controlledValues: RequestsOnly
We deployed VPA in "Off" mode first, which collects metrics and generates recommendations without changing anything. After two weeks of observation, VPA gave us target recommendations:
VPA Recommendations (after 2 weeks of observation):
─────────────────────────────────────────────────────────────────────
Service Current Request VPA Recommendation Change
─────────────────────────────────────────────────────────────────────
payments-api 1000m / 2048MB 250m / 420MB -75%
analytics-worker 1000m / 2048MB 300m / 620MB -70%
notif-service 1000m / 2048MB 320m / 510MB -68%
orders-api 1000m / 2048MB 400m / 680MB -60%
users-api 1000m / 2048MB 380m / 580MB -62%
inventory-sync 1000m / 2048MB 450m / 840MB -55%
web-bff 1000m / 2048MB 520m / 900MB -48%
api-gateway 1000m / 2048MB 600m / 1100MB -40%
─────────────────────────────────────────────────────────────────────
Cluster total req 42000m / 86GB 12800m / 28GB -70%
Step 3: Gradual Rollout
We didn't flip VPA to "Auto" mode all at once. We rolled it out namespace by namespace over two weeks, watching metrics closely between each rollout.
# vpa/vpa-auto-mode.yaml
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: payments-api-vpa
namespace: payments
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: payments-api
updatePolicy:
updateMode: "Auto"
minReplicas: 2 # Never evict below 2 replicas during resize
resourcePolicy:
containerPolicies:
- containerName: payments-api
minAllowed:
cpu: "100m"
memory: "128Mi"
maxAllowed:
cpu: "2000m"
memory: "4Gi"
controlledResources: ["cpu", "memory"]
controlledValues: RequestsOnly
We also added a safety margin by setting our HPA targets slightly higher, so pods wouldn't get squeezed between VPA shrinking requests and HPA expecting headroom.
Impact of Right-Sizing
Right-Sizing Results:
──────────────────────────────────────────────────
Total CPU requested: 42,000m → 12,800m (-70%)
Total memory requested: 86 GB → 28 GB (-67%)
Nodes required: 12 → 8 (scheduler can pack more tightly)
Monthly compute savings: ~$1,970
Zero performance impact: p99 latency unchanged
Right-sizing alone would have saved us $1,970/month. But it also unlocked the next strategies — because with realistic resource requests, the autoscaler could make much smarter decisions about how many nodes we actually needed.
Strategy 2: Cluster Autoscaler to Karpenter Migration
With right-sized pods, the Cluster Autoscaler should have been able to consolidate down to fewer nodes. But it wasn't. Nodes that dropped to 20% utilization sat there for hours. Scale-down events took 10+ minutes. And every new node was the same m5.2xlarge, even when we only needed a few hundred millicores of capacity.
Karpenter solved all of these problems.
Why Cluster Autoscaler Was Failing Us
Cluster_Autoscaler_Problems:
slow_scale_down:
description: "10-minute evaluation window before considering scale-down"
impact: "Nodes sit underutilized for 10+ minutes after load drops"
root_cause: "--scale-down-unneeded-time=10m (default)"
wrong_instance_types:
description: "Only provisions one instance type (m5.2xlarge)"
impact: "Over-provisions when only 2 vCPU needed, spins up 8 vCPU node"
root_cause: "Node groups locked to single instance type"
no_consolidation:
description: "Won't move pods between nodes to free up underutilized ones"
impact: "5 nodes at 30% utilization instead of 2 nodes at 75%"
root_cause: "Cluster Autoscaler doesn't do bin-packing consolidation"
pdb_paralysis:
description: "Won't drain node if any pod has a PodDisruptionBudget"
impact: "Nodes with a single PDB-protected pod never get removed"
root_cause: "Conservative default behavior"
Karpenter NodePool Configuration
Karpenter is fundamentally different from Cluster Autoscaler. Instead of managing fixed node groups, it provisions individual nodes based on pending pod requirements and chooses the optimal instance type dynamically.
# karpenter/nodepool-general.yaml
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
name: general-workloads
spec:
template:
metadata:
labels:
workload-type: general
spec:
requirements:
- key: kubernetes.io/arch
operator: In
values: ["amd64", "arm64"]
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand", "spot"]
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["m", "c", "r"]
- key: karpenter.k8s.aws/instance-generation
operator: Gt
values: ["5"]
- key: karpenter.k8s.aws/instance-size
operator: In
values: ["medium", "large", "xlarge", "2xlarge"]
nodeClassRef:
name: default
limits:
cpu: "100"
memory: "400Gi"
disruption:
consolidationPolicy: WhenUnderutilized
consolidateAfter: 30s
expireAfter: 720h # Replace nodes every 30 days for patches
weight: 50
# karpenter/ec2nodeclass-default.yaml
apiVersion: karpenter.k8s.aws/v1beta1
kind: EC2NodeClass
metadata:
name: default
spec:
amiFamily: AL2
subnetSelectorTerms:
- tags:
karpenter.sh/discovery: production-cluster
securityGroupSelectorTerms:
- tags:
karpenter.sh/discovery: production-cluster
instanceProfile: KarpenterNodeInstanceProfile-production
blockDeviceMappings:
- deviceName: /dev/xvda
ebs:
volumeType: gp3
volumeSize: 80Gi
deleteOnTermination: true
throughput: 125
iops: 3000
tags:
Environment: production
ManagedBy: karpenter
CostCenter: platform
metadataOptions:
httpEndpoint: enabled
httpProtocolIPv6: disabled
httpPutResponseHopLimit: 2
httpTokens: required
The Consolidation Difference
The key configuration is consolidationPolicy: WhenUnderutilized with consolidateAfter: 30s. This tells Karpenter to actively look for opportunities to consolidate workloads onto fewer nodes and terminate empty or underutilized nodes within 30 seconds. Compare that to Cluster Autoscaler's 10-minute window.
We watched the consolidation happen in real-time during our first deploy:
Karpenter Consolidation Log (first hour after migration):
──────────────────────────────────────────────────────────────────
14:02:31 Detected underutilized node ip-10-0-3-142 (22% CPU)
14:02:33 Identified replacement: consolidating pods to ip-10-0-3-98
14:02:35 Cordoning ip-10-0-3-142
14:02:38 Draining pods (3 pods, respecting PDBs)
14:02:52 All pods rescheduled successfully
14:02:54 Terminating ip-10-0-3-142
14:03:01 Node terminated. Cluster: 11 → 10 nodes
14:15:42 Detected underutilized node ip-10-0-1-87 (18% CPU)
14:15:44 Identified replacement: m5.large (vs current m5.2xlarge)
14:15:46 Launching m5.large ip-10-0-1-203 (right-sized replacement)
14:15:58 Node ready, migrating pods
14:16:14 Original node terminated. Saved: $0.192/hr → $0.096/hr
14:31:08 Detected 2 underutilized nodes, consolidating to 1
14:31:42 Consolidation complete. Cluster: 10 → 8 nodes
Within the first hour, Karpenter had already removed 4 nodes and right-sized 2 others. Over the next 48 hours, the cluster stabilized at 5 nodes on average — down from 12.
Migration Impact
Cluster Autoscaler → Karpenter Migration:
──────────────────────────────────────────────────────────
Average node count: 12 → 5 nodes
Scale-down time: 10+ minutes → 30 seconds
Instance type variety: 1 type → 15+ types (auto-selected)
Node utilization: 38% → 68% (before other optimizations)
Monthly compute savings: ~$3,450
Strategy 3: Spot Instances for Non-Critical Workloads
With Karpenter handling node provisioning, adding spot instances was straightforward. The key was being deliberate about which workloads could tolerate interruption and which couldn't.
Workload Classification
We classified every workload into three tiers:
Workload_Tiers:
tier_1_critical:
description: "Customer-facing APIs, payment processing"
capacity_type: "on-demand only"
examples:
- api-gateway
- payments-api
- orders-api
percentage_of_cluster: 35%
tier_2_important:
description: "Internal services, async processors"
capacity_type: "on-demand preferred, spot acceptable"
examples:
- users-api
- inventory-sync
- notification-service
percentage_of_cluster: 30%
tier_3_flexible:
description: "Batch jobs, analytics, dev/staging"
capacity_type: "spot preferred"
examples:
- analytics-worker
- report-generator
- staging-environments
- cron-jobs
percentage_of_cluster: 35%
Karpenter NodePool for Spot Workloads
# karpenter/nodepool-spot.yaml
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
name: spot-workloads
spec:
template:
metadata:
labels:
workload-type: spot-eligible
capacity-type: spot
spec:
requirements:
- key: kubernetes.io/arch
operator: In
values: ["amd64", "arm64"]
- key: karpenter.sh/capacity-type
operator: In
values: ["spot"]
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["m", "c", "r"]
- key: karpenter.k8s.aws/instance-generation
operator: Gt
values: ["5"]
- key: karpenter.k8s.aws/instance-size
operator: In
values: ["large", "xlarge", "2xlarge"]
nodeClassRef:
name: default
# Taint so only spot-tolerant pods land here
taints:
- key: capacity-type
value: spot
effect: NoSchedule
limits:
cpu: "60"
memory: "240Gi"
disruption:
consolidationPolicy: WhenUnderutilized
consolidateAfter: 15s
expireAfter: 168h # Replace weekly — spot nodes are ephemeral
weight: 80 # Higher weight = Karpenter prefers this pool
Pod Configuration for Spot Tolerance
# deployments/analytics-worker.yaml (Tier 3 — spot eligible)
apiVersion: apps/v1
kind: Deployment
metadata:
name: analytics-worker
namespace: analytics
spec:
replicas: 3
selector:
matchLabels:
app: analytics-worker
template:
metadata:
labels:
app: analytics-worker
tier: flexible
spec:
# Tolerate spot taint — allows scheduling on spot nodes
tolerations:
- key: capacity-type
value: spot
operator: Equal
effect: NoSchedule
# Prefer spot nodes, but allow on-demand if spot unavailable
affinity:
nodeAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 90
preference:
matchExpressions:
- key: karpenter.sh/capacity-type
operator: In
values: ["spot"]
- weight: 10
preference:
matchExpressions:
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand"]
# Spread across AZs for spot diversity
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfied: DoNotSchedule
labelSelector:
matchLabels:
app: analytics-worker
# Graceful shutdown on spot interruption
terminationGracePeriodSeconds: 120
containers:
- name: analytics-worker
image: 123456789.dkr.ecr.us-east-1.amazonaws.com/analytics-worker:v2.4.1
resources:
requests:
cpu: "300m"
memory: "620Mi"
limits:
cpu: "600m"
memory: "1Gi"
# Handle SIGTERM gracefully for spot interruptions
lifecycle:
preStop:
exec:
command:
- /bin/sh
- -c
- |
echo "Spot interruption: draining in-flight jobs..."
curl -s -X POST http://localhost:8080/admin/drain
sleep 30
Pod Disruption Budgets for Spot Safety
# pdb/analytics-worker-pdb.yaml
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: analytics-worker-pdb
namespace: analytics
spec:
maxUnavailable: 1 # Allow 1 pod at a time to be disrupted
selector:
matchLabels:
app: analytics-worker
Spot Instance Savings
Spot Instance Breakdown:
──────────────────────────────────────────────────────────────
Workload Tier Pods On-Demand Spot Spot Savings
──────────────────────────────────────────────────────────────
Tier 1 (critical) 14 14 (100%) 0 (0%) $0
Tier 2 (important) 12 4 (33%) 8 (67%) $380/mo
Tier 3 (flexible) 16 2 (12%) 14 (88%) $720/mo
──────────────────────────────────────────────────────────────
Total 42 20 (48%) 22 (52%) $1,100/mo
Spot interruption rate (30-day average): 3.2%
Interruptions handled gracefully: 100%
Customer impact from interruptions: Zero
In practice, about 65% of our actual compute hours ended up running on spot (since Tier 3 workloads consumed more CPU/memory proportionally). Spot pricing averaged 60-70% cheaper than on-demand for the instance types Karpenter selected.
Strategy 4: Namespace Resource Quotas and Cost Allocation
This was the strategy that surprised us most. Not because of technical complexity — it was straightforward — but because of how dramatically team behavior changed when they could see their own costs.
Resource Quotas per Namespace
# quotas/payments-namespace-quota.yaml
apiVersion: v1
kind: ResourceQuota
metadata:
name: payments-resource-quota
namespace: payments
spec:
hard:
requests.cpu: "4000m" # 4 vCPU total for namespace
requests.memory: "8Gi" # 8 GB RAM total
limits.cpu: "8000m"
limits.memory: "16Gi"
pods: "20" # Max 20 pods
persistentvolumeclaims: "5"
services.loadbalancers: "1"
# quotas/payments-limit-range.yaml
apiVersion: v1
kind: LimitRange
metadata:
name: payments-limit-range
namespace: payments
spec:
limits:
# Default limits applied to containers that don't specify their own
- type: Container
default:
cpu: "500m"
memory: "512Mi"
defaultRequest:
cpu: "200m"
memory: "256Mi"
min:
cpu: "50m"
memory: "64Mi"
max:
cpu: "2000m"
memory: "4Gi"
# Pod-level limits
- type: Pod
max:
cpu: "4000m"
memory: "8Gi"
Cost Allocation Labels
We enforced a labeling standard across all deployments. Karpenter propagated these labels to nodes, and we used them for cost allocation in our reporting.
# Standard labels required on all deployments
metadata:
labels:
app.kubernetes.io/name: payments-api
app.kubernetes.io/component: api
team: payments # Cost allocation
environment: production # Environment tracking
cost-center: engineering-payments # Finance mapping
tier: critical # Workload classification
We enforced the labeling requirement with an OPA Gatekeeper policy:
# policies/require-cost-labels.yaml
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredLabels
metadata:
name: require-cost-labels
spec:
match:
kinds:
- apiGroups: ["apps"]
kinds: ["Deployment", "StatefulSet", "DaemonSet"]
excludedNamespaces: ["kube-system", "karpenter", "monitoring"]
parameters:
labels:
- key: team
allowedRegex: "^[a-z]([a-z0-9-]*[a-z0-9])?$"
- key: cost-center
- key: tier
allowedRegex: "^(critical|important|flexible)$"
Per-Team Cost Reporting Script
#!/usr/bin/env python3
"""
cost_report.py — Generate per-team Kubernetes cost allocation reports.
Pulls pod resource usage, maps to node costs, and allocates by namespace/team.
"""
import subprocess
import json
import boto3
from datetime import datetime, timedelta
from collections import defaultdict
# Approximate hourly cost per vCPU-hour and GB-hour (blended on-demand + spot)
CPU_COST_PER_HOUR = 0.0425 # ~$0.0425 per vCPU-hour (blended rate)
MEM_COST_PER_GB_HOUR = 0.005 # ~$0.005 per GB-hour
HOURS_IN_MONTH = 730
def get_namespace_usage() -> dict:
"""Get resource usage aggregated by namespace and team label."""
result = subprocess.run(
["kubectl", "get", "pods", "--all-namespaces",
"-o", "json"],
capture_output=True, text=True, check=True
)
pods = json.loads(result.stdout)
# Get current metrics
metrics_result = subprocess.run(
["kubectl", "top", "pods", "--all-namespaces",
"--no-headers"],
capture_output=True, text=True, check=True
)
# Parse metrics output
usage_by_pod = {}
for line in metrics_result.stdout.strip().split("\n"):
parts = line.split()
if len(parts) >= 4:
ns, pod, cpu, mem = parts[0], parts[1], parts[2], parts[3]
usage_by_pod[f"{ns}/{pod}"] = {
"cpu_m": int(cpu.replace("m", "")),
"mem_mb": int(mem.replace("Mi", ""))
}
# Aggregate by team
team_costs = defaultdict(lambda: {
"cpu_requested_m": 0,
"cpu_used_m": 0,
"mem_requested_mb": 0,
"mem_used_mb": 0,
"pod_count": 0,
"namespaces": set()
})
for pod in pods.get("items", []):
ns = pod["metadata"]["namespace"]
if ns in ["kube-system", "karpenter", "monitoring", "kube-node-lease"]:
continue
labels = pod["metadata"].get("labels", {})
team = labels.get("team", "untagged")
pod_key = f"{ns}/{pod['metadata']['name']}"
team_costs[team]["namespaces"].add(ns)
team_costs[team]["pod_count"] += 1
for container in pod["spec"].get("containers", []):
requests = container.get("resources", {}).get("requests", {})
cpu_req = requests.get("cpu", "0m")
mem_req = requests.get("memory", "0Mi")
# Parse CPU
if cpu_req.endswith("m"):
team_costs[team]["cpu_requested_m"] += int(cpu_req[:-1])
else:
team_costs[team]["cpu_requested_m"] += int(float(cpu_req) * 1000)
# Parse memory
if mem_req.endswith("Mi"):
team_costs[team]["mem_requested_mb"] += int(mem_req[:-2])
elif mem_req.endswith("Gi"):
team_costs[team]["mem_requested_mb"] += int(float(mem_req[:-2]) * 1024)
# Add actual usage
usage = usage_by_pod.get(pod_key, {"cpu_m": 0, "mem_mb": 0})
team_costs[team]["cpu_used_m"] += usage["cpu_m"]
team_costs[team]["mem_used_mb"] += usage["mem_mb"]
return team_costs
def calculate_costs(team_costs: dict) -> list[dict]:
"""Calculate estimated monthly costs per team."""
reports = []
for team, data in team_costs.items():
cpu_vcpu = data["cpu_requested_m"] / 1000
mem_gb = data["mem_requested_mb"] / 1024
monthly_cpu_cost = cpu_vcpu * CPU_COST_PER_HOUR * HOURS_IN_MONTH
monthly_mem_cost = mem_gb * MEM_COST_PER_GB_HOUR * HOURS_IN_MONTH
# Waste calculation
cpu_waste_pct = (
(data["cpu_requested_m"] - data["cpu_used_m"])
/ data["cpu_requested_m"] * 100
if data["cpu_requested_m"] > 0 else 0
)
reports.append({
"team": team,
"namespaces": sorted(data["namespaces"]),
"pods": data["pod_count"],
"cpu_requested": f"{cpu_vcpu:.1f} vCPU",
"mem_requested": f"{mem_gb:.1f} GB",
"monthly_cost": round(monthly_cpu_cost + monthly_mem_cost, 2),
"cpu_waste_pct": round(cpu_waste_pct, 1),
"potential_savings": round(
(monthly_cpu_cost + monthly_mem_cost) * cpu_waste_pct / 100, 2
),
})
reports.sort(key=lambda r: r["monthly_cost"], reverse=True)
return reports
def print_cost_report(reports: list[dict]):
"""Print formatted cost allocation report."""
print(f"\n{'='*80}")
print(f" EKS Cost Allocation Report — {datetime.now().strftime('%B %Y')}")
print(f"{'='*80}\n")
total_cost = 0
total_savings = 0
print(f"{'Team':<18} {'Pods':>5} {'CPU':>8} {'Memory':>8} "
f"{'Monthly':>10} {'Waste%':>7} {'Saveable':>10}")
print("─" * 80)
for r in reports:
total_cost += r["monthly_cost"]
total_savings += r["potential_savings"]
print(f"{r['team']:<18} {r['pods']:>5} {r['cpu_requested']:>8} "
f"{r['mem_requested']:>8} ${r['monthly_cost']:>8,.2f} "
f"{r['cpu_waste_pct']:>6.1f}% ${r['potential_savings']:>8,.2f}")
print("─" * 80)
print(f"{'TOTAL':<18} {'':>5} {'':>8} {'':>8} ${total_cost:>8,.2f} "
f"{'':>7} ${total_savings:>8,.2f}")
print(f"\nTotal monthly compute cost: ${total_cost:,.2f}")
print(f"Identified savings potential: ${total_savings:,.2f} "
f"({total_savings/total_cost*100:.0f}%)")
def send_to_cloudwatch(reports: list[dict]):
"""Publish per-team cost metrics to CloudWatch for dashboards."""
cloudwatch = boto3.client("cloudwatch")
metric_data = []
for r in reports:
metric_data.extend([
{
"MetricName": "TeamMonthlyCost",
"Value": r["monthly_cost"],
"Unit": "None",
"Dimensions": [{"Name": "Team", "Value": r["team"]}]
},
{
"MetricName": "TeamWastePercentage",
"Value": r["cpu_waste_pct"],
"Unit": "Percent",
"Dimensions": [{"Name": "Team", "Value": r["team"]}]
}
])
# CloudWatch accepts max 20 metrics per call
for i in range(0, len(metric_data), 20):
cloudwatch.put_metric_data(
Namespace="Kubernetes/CostAllocation",
MetricData=metric_data[i:i+20]
)
if __name__ == "__main__":
team_costs = get_namespace_usage()
reports = calculate_costs(team_costs)
print_cost_report(reports)
send_to_cloudwatch(reports)
The Behavioral Impact
We ran this report weekly and sent it to every engineering lead. Within one month:
Cost Allocation Impact (30 days after enabling reports):
──────────────────────────────────────────────────────────────
Team Before Report After Report Reduction
──────────────────────────────────────────────────────────────
payments $820/mo $490/mo -40%
analytics $680/mo $340/mo -50%
orders $720/mo $460/mo -36%
users $540/mo $340/mo -37%
notifications $480/mo $290/mo -40%
inventory $620/mo $380/mo -39%
frontend $440/mo $310/mo -30%
platform $380/mo $280/mo -26%
──────────────────────────────────────────────────────────────
Total $4,680/mo $2,890/mo -38%
Teams voluntarily reduced their resource requests, removed unused deployments, and consolidated microservices once they could see the dollar amounts. The payments team discovered they were running 3 replicas of a debug service that hadn't been used in 6 months. The analytics team moved their batch jobs to spot instances. Nobody asked them to — they just started caring when the numbers were visible.
Strategy 5: Node Consolidation and Bin-Packing
Even after right-sizing and switching to Karpenter, we found that pods weren't being distributed optimally. Some nodes were packed at 85% while others sat at 40%, because the default Kubernetes scheduler spreads pods across nodes for high availability without considering bin-packing efficiency.
Karpenter Consolidation Policy (Enhanced)
We already had consolidationPolicy: WhenUnderutilized, but we tuned it further:
# karpenter/nodepool-consolidated.yaml
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
name: consolidated-general
spec:
template:
spec:
requirements:
- key: kubernetes.io/arch
operator: In
values: ["amd64", "arm64"]
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand", "spot"]
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["m", "c", "r"]
- key: karpenter.k8s.aws/instance-generation
operator: Gt
values: ["5"]
nodeClassRef:
name: default
disruption:
consolidationPolicy: WhenUnderutilized
consolidateAfter: 30s
# Budget controls how many nodes can be disrupted simultaneously
budgets:
- nodes: "20%" # Max 20% of nodes disrupted at once
- nodes: "0" # No disruptions during business hours peak
schedule: "0 9 * * 1-5" # Mon-Fri 9 AM
duration: 2h # For 2 hours
Pod Topology Spread Constraints
We wanted pods spread across AZs for availability, but packed tightly within each AZ for efficiency. Topology spread constraints gave us both.
# Example deployment with topology spread
apiVersion: apps/v1
kind: Deployment
metadata:
name: orders-api
namespace: orders
spec:
replicas: 4
selector:
matchLabels:
app: orders-api
template:
metadata:
labels:
app: orders-api
team: orders
tier: critical
spec:
topologySpreadConstraints:
# Spread across AZs — hard requirement for HA
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfied: DoNotSchedule
labelSelector:
matchLabels:
app: orders-api
# Pack within each node — soft preference for bin-packing
- maxSkew: 2
topologyKey: kubernetes.io/hostname
whenUnsatisfied: ScheduleAnyway
labelSelector:
matchLabels:
app: orders-api
containers:
- name: orders-api
image: 123456789.dkr.ecr.us-east-1.amazonaws.com/orders-api:v3.1.0
resources:
requests:
cpu: "400m"
memory: "680Mi"
limits:
cpu: "800m"
memory: "1Gi"
Descheduler for Rebalancing
Over time, pods accumulate on nodes unevenly — especially after spot interruptions or rolling deployments. The Kubernetes Descheduler periodically evicts pods from underutilized or imbalanced nodes so the scheduler can re-place them optimally.
# descheduler/config.yaml
apiVersion: descheduler/v1alpha2
kind: DeschedulerPolicy
profiles:
- name: cost-optimization
pluginConfig:
- name: RemoveDuplicates
args:
excludeOwnerKinds:
- DaemonSet
namespaces:
exclude:
- kube-system
- monitoring
- name: LowNodeUtilization
args:
thresholds:
cpu: 30
memory: 30
pods: 20
targetThresholds:
cpu: 60
memory: 60
pods: 40
numberOfNodes: 2 # Only rebalance if 2+ nodes are underutilized
- name: RemovePodsHavingTooManyRestarts
args:
podRestartThreshold: 10
includingInitContainers: true
plugins:
balance:
enabled:
- RemoveDuplicates
- LowNodeUtilization
deschedule:
enabled:
- RemovePodsHavingTooManyRestarts
# descheduler/cronjob.yaml
apiVersion: batch/v1
kind: CronJob
metadata:
name: descheduler
namespace: kube-system
spec:
schedule: "*/30 * * * *" # Every 30 minutes
concurrencyPolicy: Forbid
jobTemplate:
spec:
template:
spec:
serviceAccountName: descheduler
containers:
- name: descheduler
image: registry.k8s.io/descheduler/descheduler:v0.28.0
args:
- --policy-config-file=/policy/config.yaml
- --v=3
volumeMounts:
- name: policy
mountPath: /policy
volumes:
- name: policy
configMap:
name: descheduler-policy
restartPolicy: Never
Consolidation Results
Bin-Packing / Consolidation Results:
──────────────────────────────────────────────────────────
Metric Before After Change
──────────────────────────────────────────────────────────
Avg node CPU utilization 52% 72% +38%
Avg node mem utilization 48% 68% +42%
Nodes with <30% CPU 3-4 0-1 -80%
Average node count 7 5 -29%
Monthly compute savings — $980/mo —
Strategy 6: Graviton/ARM Instances
This was our final optimization — and one of the easiest to implement with Karpenter already in place. AWS Graviton3 (ARM-based) instances are approximately 20% cheaper than equivalent x86 instances and deliver better performance for most workloads.
Multi-Architecture Docker Builds
The prerequisite was making our container images work on both amd64 and arm64. We updated our CI pipeline to build multi-arch images:
# .github/workflows/build-multiarch.yaml
name: Build Multi-Arch Image
on:
push:
branches: [main]
jobs:
build:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Set up QEMU for multi-arch builds
uses: docker/setup-qemu-action@v3
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Login to ECR
uses: aws-actions/amazon-ecr-login@v2
- name: Build and push multi-arch image
uses: docker/build-push-action@v5
with:
context: .
platforms: linux/amd64,linux/arm64
push: true
tags: |
123456789.dkr.ecr.us-east-1.amazonaws.com/orders-api:${{ github.sha }}
123456789.dkr.ecr.us-east-1.amazonaws.com/orders-api:latest
cache-from: type=gha
cache-to: type=gha,mode=max
Karpenter NodePool Preferring Graviton
# karpenter/nodepool-graviton.yaml
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
name: graviton-preferred
spec:
template:
metadata:
labels:
workload-type: general
architecture: multi-arch
spec:
requirements:
# Prefer ARM (Graviton), but allow x86 as fallback
- key: kubernetes.io/arch
operator: In
values: ["arm64", "amd64"]
- key: karpenter.sh/capacity-type
operator: In
values: ["on-demand", "spot"]
- key: karpenter.k8s.aws/instance-category
operator: In
values: ["m", "c", "r"]
- key: karpenter.k8s.aws/instance-generation
operator: Gt
values: ["6"] # Graviton3+ only
nodeClassRef:
name: default
# Lower weight = higher priority
# Karpenter will prefer cheaper Graviton instances automatically
weight: 10
disruption:
consolidationPolicy: WhenUnderutilized
consolidateAfter: 30s
Karpenter's cost-aware scheduling naturally prefers Graviton instances because they're cheaper. We didn't need to add explicit affinity rules — Karpenter calculates the cheapest instance type that fits the pending pods and picks Graviton when it's the best option. Within a week, 60% of our nodes were Graviton.
Graviton Performance Comparison
We ran our standard load tests on both architectures:
Graviton3 vs x86 (m7g.xlarge vs m5.xlarge):
────────────────────────────────────────────────────────────────
Metric x86 (m5.xlarge) Graviton (m7g.xlarge) Diff
────────────────────────────────────────────────────────────────
On-demand price/hr $0.192 $0.1632 -15%
Spot price/hr (avg) $0.072 $0.058 -19%
API p50 latency 12ms 10ms -17%
API p99 latency 45ms 38ms -16%
Requests/sec (max) 8,200 9,400 +15%
Container startup 520ms 480ms -8%
────────────────────────────────────────────────────────────────
Graviton was cheaper AND faster. The 15% price reduction on on-demand and 19% on spot pricing compounded with our other savings.
Architecture Migration Impact
Graviton Migration Results:
──────────────────────────────────────────────────────────
Nodes running Graviton: 60% (3 of 5 avg)
On-demand cost reduction: -20% per node-hour
Spot cost reduction: -19% per node-hour
Performance improvement: 15% higher throughput
Monthly compute savings: $520/mo
Build pipeline overhead: +3 min (multi-arch build)
The only catch was the multi-arch Docker build, which added about 3 minutes to our CI pipeline. We offset this with BuildKit layer caching — the arm64 layers are cached in GitHub Actions, so subsequent builds only rebuild changed layers.
Monitoring: Cost Visibility That Actually Drives Change
We needed ongoing monitoring, not just a one-time audit. We deployed Kubecost alongside custom CloudWatch metrics to create a cost observability layer.
Kubecost Deployment
# kubecost/values.yaml (Helm)
kubecostProductConfigs:
clusterName: production-cluster
clusterProfile: production
currencyCode: USD
customPricesEnabled: true
# Use AWS spot pricing data for accurate cost calculation
kubecostModel:
etlCloudAsset: true
# CloudWatch integration for custom dashboards
cloudIntegrationSecret:
enabled: true
data:
cloud-integration.json: |
{
"aws": {
"athenaBucketName": "s3://kubecost-athena-results",
"athenaRegion": "us-east-1",
"athenaDatabase": "athenacurcfn_kube_costs",
"athenaTable": "kube_costs",
"masterPayerARN": "",
"serviceKeyName": ""
}
}
# Resource settings for Kubecost itself
prometheus:
server:
resources:
requests:
cpu: "200m"
memory: "512Mi"
retention: "15d"
Custom CloudWatch Metrics for Cluster Efficiency
#!/usr/bin/env python3
"""
cluster_efficiency_metrics.py — Publish daily cluster efficiency metrics
to CloudWatch for dashboards and alerting.
Run as a CronJob in the cluster: schedule "0 8 * * *" (daily at 8 AM)
"""
import subprocess
import json
import boto3
from datetime import datetime
def get_cluster_metrics() -> dict:
"""Gather comprehensive cluster efficiency metrics."""
# Node metrics
nodes_result = subprocess.run(
["kubectl", "get", "nodes", "-o", "json"],
capture_output=True, text=True, check=True
)
nodes = json.loads(nodes_result.stdout)["items"]
# Node resource usage
top_result = subprocess.run(
["kubectl", "top", "nodes", "--no-headers"],
capture_output=True, text=True, check=True
)
node_metrics = []
for line in top_result.stdout.strip().split("\n"):
parts = line.split()
if len(parts) >= 5:
node_metrics.append({
"name": parts[0],
"cpu_m": int(parts[1].replace("m", "")),
"cpu_pct": int(parts[2].replace("%", "")),
"mem_mb": int(parts[3].replace("Mi", "")),
"mem_pct": int(parts[4].replace("%", ""))
})
# Pod counts
pods_result = subprocess.run(
["kubectl", "get", "pods", "--all-namespaces",
"--field-selector=status.phase=Running", "--no-headers"],
capture_output=True, text=True, check=True
)
running_pods = len(pods_result.stdout.strip().split("\n"))
# Spot vs on-demand node counts
spot_nodes = 0
od_nodes = 0
graviton_nodes = 0
for node in nodes:
labels = node["metadata"].get("labels", {})
if labels.get("karpenter.sh/capacity-type") == "spot":
spot_nodes += 1
else:
od_nodes += 1
if labels.get("kubernetes.io/arch") == "arm64":
graviton_nodes += 1
total_nodes = len(nodes)
avg_cpu = sum(n["cpu_pct"] for n in node_metrics) / len(node_metrics) if node_metrics else 0
avg_mem = sum(n["mem_pct"] for n in node_metrics) / len(node_metrics) if node_metrics else 0
return {
"total_nodes": total_nodes,
"spot_nodes": spot_nodes,
"on_demand_nodes": od_nodes,
"graviton_nodes": graviton_nodes,
"avg_cpu_utilization": round(avg_cpu, 1),
"avg_mem_utilization": round(avg_mem, 1),
"running_pods": running_pods,
"spot_percentage": round(spot_nodes / total_nodes * 100, 1) if total_nodes > 0 else 0,
"graviton_percentage": round(graviton_nodes / total_nodes * 100, 1) if total_nodes > 0 else 0,
}
def publish_metrics(metrics: dict):
"""Publish cluster metrics to CloudWatch."""
cloudwatch = boto3.client("cloudwatch")
metric_data = [
{
"MetricName": "TotalNodes",
"Value": metrics["total_nodes"],
"Unit": "Count",
},
{
"MetricName": "SpotPercentage",
"Value": metrics["spot_percentage"],
"Unit": "Percent",
},
{
"MetricName": "GravitonPercentage",
"Value": metrics["graviton_percentage"],
"Unit": "Percent",
},
{
"MetricName": "AvgCPUUtilization",
"Value": metrics["avg_cpu_utilization"],
"Unit": "Percent",
},
{
"MetricName": "AvgMemoryUtilization",
"Value": metrics["avg_mem_utilization"],
"Unit": "Percent",
},
{
"MetricName": "RunningPods",
"Value": metrics["running_pods"],
"Unit": "Count",
},
]
# Add cluster name dimension to all metrics
for m in metric_data:
m["Dimensions"] = [
{"Name": "ClusterName", "Value": "production-cluster"}
]
cloudwatch.put_metric_data(
Namespace="Kubernetes/ClusterEfficiency",
MetricData=metric_data
)
print(f"Published {len(metric_data)} metrics to CloudWatch")
def print_daily_summary(metrics: dict):
"""Print a daily efficiency summary."""
print(f"\n{'='*60}")
print(f" Cluster Efficiency Report — {datetime.now().strftime('%Y-%m-%d')}")
print(f"{'='*60}")
print(f" Nodes: {metrics['total_nodes']} "
f"({metrics['spot_nodes']} spot, {metrics['on_demand_nodes']} on-demand)")
print(f" Graviton: {metrics['graviton_nodes']} nodes "
f"({metrics['graviton_percentage']}%)")
print(f" CPU util: {metrics['avg_cpu_utilization']}%")
print(f" Memory util: {metrics['avg_mem_utilization']}%")
print(f" Running pods: {metrics['running_pods']}")
print(f" Spot %: {metrics['spot_percentage']}%")
# Efficiency score (simple weighted average)
score = (
metrics["avg_cpu_utilization"] * 0.3
+ metrics["avg_mem_utilization"] * 0.2
+ metrics["spot_percentage"] * 0.25
+ metrics["graviton_percentage"] * 0.15
+ min(metrics["avg_cpu_utilization"], 80) / 80 * 10 # Bonus for hitting 80%+ util
)
print(f"\n Efficiency Score: {score:.0f}/100")
if score >= 70:
print(f" Status: HEALTHY")
elif score >= 50:
print(f" Status: NEEDS ATTENTION")
else:
print(f" Status: ACTION REQUIRED")
print(f"{'='*60}\n")
if __name__ == "__main__":
metrics = get_cluster_metrics()
print_daily_summary(metrics)
publish_metrics(metrics)
CloudWatch Alerting
# cloudwatch/cluster-efficiency-alarms.yaml (CloudFormation)
AWSTemplateFormatVersion: '2010-09-09'
Description: Kubernetes cluster efficiency alarms
Resources:
LowUtilizationAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: eks-low-cpu-utilization
AlarmDescription: >
Cluster CPU utilization below 50% for 2 hours —
nodes may need consolidation
MetricName: AvgCPUUtilization
Namespace: Kubernetes/ClusterEfficiency
Statistic: Average
Period: 3600
EvaluationPeriods: 2
Threshold: 50
ComparisonOperator: LessThanThreshold
Dimensions:
- Name: ClusterName
Value: production-cluster
AlarmActions:
- !Ref OpsAlertTopic
HighUtilizationAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: eks-high-cpu-utilization
AlarmDescription: >
Cluster CPU utilization above 85% for 15 minutes —
may need more capacity
MetricName: AvgCPUUtilization
Namespace: Kubernetes/ClusterEfficiency
Statistic: Average
Period: 300
EvaluationPeriods: 3
Threshold: 85
ComparisonOperator: GreaterThanThreshold
Dimensions:
- Name: ClusterName
Value: production-cluster
AlarmActions:
- !Ref OpsAlertTopic
LowSpotPercentageAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: eks-low-spot-percentage
AlarmDescription: >
Spot instance percentage dropped below 40% —
check spot availability or pricing
MetricName: SpotPercentage
Namespace: Kubernetes/ClusterEfficiency
Statistic: Average
Period: 3600
EvaluationPeriods: 1
Threshold: 40
ComparisonOperator: LessThanThreshold
Dimensions:
- Name: ClusterName
Value: production-cluster
AlarmActions:
- !Ref OpsAlertTopic
OpsAlertTopic:
Type: AWS::SNS::Topic
Properties:
TopicName: eks-cost-optimization-alerts
Subscription:
- Protocol: email
Endpoint: platform-team@company.com
Results: Before vs After
Comprehensive Results (6-week optimization):
────────────────────────────────────────────────────────────────────────────
Metric Before After Change
────────────────────────────────────────────────────────────────────────────
Monthly EKS spend $14,200 $4,180 -71%
Annual spend $170,400 $50,160 -71%
Node count (average) 12 5 -58%
Instance types in use 1 15+ Dynamic
CPU utilization 38% 72% +89%
Memory utilization 42% 68% +62%
Spot instance percentage 0% 65% +65%
Graviton/ARM percentage 0% 60% +60%
Pod CPU waste (req vs used) 73% 18% -75%
Scale-down latency 10+ min 30 sec -95%
Cost allocation visibility None Per-team Full
P99 API latency 42ms 36ms -14%
Availability 99.95% 99.97% +0.02%
Production incidents 0 0 Zero
────────────────────────────────────────────────────────────────────────────
Savings Breakdown by Strategy
Monthly Savings Attribution:
────────────────────────────────────────────────────────────────
Strategy Savings % of Total
────────────────────────────────────────────────────────────────
Strategy 1: Pod right-sizing (VPA) $1,970 19.7%
Strategy 2: Karpenter migration $3,450 34.4%
Strategy 3: Spot instances $1,100 11.0%
Strategy 4: Cost allocation (behavior) $1,790 17.9%
Strategy 5: Bin-packing/consolidation $980 9.8%
Strategy 6: Graviton instances $520 5.2%
Other (NAT, data transfer, EBS) $210 2.1%
────────────────────────────────────────────────────────────────
Total Monthly Savings $10,020 100%
Annual Savings $120,240
────────────────────────────────────────────────────────────────
ROI Analysis
Investment:
- Engineering time: 2 engineers × 6 weeks (part-time) $18,000
- Kubecost license (open source core): $0
- Karpenter (open source): $0
- Multi-arch CI pipeline updates: $500
- Load testing and validation: $1,200
- Documentation and runbooks: $800
Total Investment: $20,500
Returns:
- Monthly infrastructure savings: $10,020
- Reduced on-call toil (fewer scaling issues): $400/month
- Faster deployments (smaller images, faster scaling): $300/month
Total Monthly Savings: $10,720
Payback Period: 1.9 months
Annual Net Savings: $128,640 - $20,500 = $108,140
3-Year Net Savings: $365,420
ROI: 528% (first year)
The payback period under two months makes this one of the highest-ROI infrastructure projects we've done. And unlike a one-time optimization, these savings compound — as we add more services, they automatically get right-sized by VPA, scheduled onto spot/Graviton by Karpenter, and tracked by cost allocation dashboards.
Lessons Learned
What Delivered the Biggest Impact
Karpenter over Cluster Autoscaler was the single biggest win. The combination of intelligent instance selection, fast consolidation, and native spot support accounted for 34% of our total savings. If you only do one thing from this post, migrate to Karpenter.
Cost visibility changed team behavior more than any technical optimization. When engineering leads could see that their namespace was costing $820/month with 75% waste, they fixed it themselves. We didn't write a single Jira ticket — they just started right-sizing their own pods. This alone drove 18% of our savings.
VPA in "Off" mode for 2 weeks was critical. We were tempted to skip straight to "Auto" mode. Glad we didn't. The recommendation-only period let us catch two services where VPA's recommendation was too aggressive (a bursty cron job and an ML inference service with spiky memory usage). We set custom min/max bounds for those before enabling auto-scaling.
Pod right-sizing unlocked everything else. Until pods had realistic resource requests, the scheduler and autoscaler were making decisions based on fiction. Right-sizing was the foundation that made every other strategy effective.
Mistakes We Made
We forgot about PodDisruptionBudgets during the Karpenter migration. Several PDBs had
minAvailableset equal to the replica count, which meant Karpenter couldn't evict any pods for consolidation. We had to audit every PDB and fix the ones that were too restrictive. Lesson: audit PDBs before enabling consolidation.Our first Graviton rollout broke one service. One team had a Python dependency (
cryptography38.x) that didn't have pre-built ARM64 wheels. The build worked (it compiled from source), but it added 8 minutes to their Docker build. We pinned that service to amd64 until they upgraded the dependency. Lesson: test multi-arch builds per-service before enforcing cluster-wide.We set Karpenter consolidation too aggressively at first. With
consolidateAfter: 0s, Karpenter was constantly shuffling pods during traffic fluctuations. We bumped it to 30 seconds, which smoothed things out without sacrificing efficiency. Lesson: consolidation needs a buffer to avoid thrashing.Spot interruptions during a deploy caused a 2-minute partial outage. We didn't have enough on-demand capacity for our Tier 1 services when a spot reclamation happened during a rolling update. We tightened our tier classifications and ensured critical services never touch spot nodes. Lesson: classify workloads before enabling spot, not after.
What Surprised Us
Graviton was faster, not just cheaper. We expected a cost reduction and assumed equivalent performance. Instead, our API latency dropped 14% on Graviton nodes. The ARM architecture handles our Python/Node workloads more efficiently than we expected.
The descheduler had an outsized impact on bin-packing. Even with Karpenter's consolidation, pod distribution drifted over time. Running the descheduler every 30 minutes consistently recovered 10-15% of wasted node capacity that Karpenter alone missed.
Teams asked for MORE cost visibility, not less. We expected pushback when we started publishing per-team costs. Instead, multiple teams asked for more granular breakdowns — per-service, per-environment, with trend lines. Cost awareness became a point of pride, not shame.
Action Items Checklist
If your EKS cluster is running with default configurations, here's the prioritized order of operations:
-
Audit Your Waste (Week 1)
- [ ] Run a resource audit comparing pod requests vs actual usage
- [ ] Identify your top 5 most over-provisioned services
- [ ] Measure current cluster CPU and memory utilization
- [ ] Document current monthly spend baseline
- [ ] Calculate per-namespace cost estimates
-
Right-Size Pods (Week 2-3)
- [ ] Deploy VPA in "Off" (recommendation) mode
- [ ] Collect 2 weeks of recommendations
- [ ] Review recommendations for bursty or spiky workloads
- [ ] Set appropriate min/max bounds per service
- [ ] Roll out VPA "Auto" mode namespace by namespace
-
Migrate to Karpenter (Week 3-4)
- [ ] Install Karpenter alongside Cluster Autoscaler
- [ ] Create NodePool with diverse instance types
- [ ] Enable consolidation with 30-second buffer
- [ ] Audit and fix PodDisruptionBudgets
- [ ] Drain managed node groups one at a time
- [ ] Remove Cluster Autoscaler after validation
-
Enable Spot Instances (Week 4-5)
- [ ] Classify workloads into critical / important / flexible tiers
- [ ] Create spot-specific NodePool with taints
- [ ] Add tolerations and affinity to flexible-tier deployments
- [ ] Configure PodDisruptionBudgets for spot workloads
- [ ] Add graceful shutdown handlers for spot interruptions
-
Implement Cost Allocation (Week 5)
- [ ] Define required labels standard (team, cost-center, tier)
- [ ] Deploy OPA Gatekeeper policy to enforce labels
- [ ] Set ResourceQuota and LimitRange per namespace
- [ ] Deploy cost reporting script as a weekly CronJob
- [ ] Share first cost report with engineering leads
-
Optimize Further (Week 5-6)
- [ ] Build multi-arch Docker images
- [ ] Enable Graviton in Karpenter NodePool
- [ ] Deploy descheduler for ongoing rebalancing
- [ ] Set up CloudWatch alarms for efficiency metrics
- [ ] Deploy Kubecost for real-time cost dashboards
-
Sustain the Gains (Ongoing)
- [ ] Review weekly cost reports and act on regressions
- [ ] Run monthly resource audits
- [ ] Update VPA bounds when services change significantly
- [ ] Review spot interruption rates quarterly
- [ ] Track efficiency score in team dashboards
Conclusion
Kubernetes cost optimization isn't a single configuration change — it's a system of practices that compound. Pod right-sizing makes autoscaling smarter. Smarter autoscaling makes spot instances viable. Cost visibility makes teams care. And Graviton makes everything cheaper.
We went from $14,200/month to $4,180/month — a 71% reduction — by applying six strategies in sequence over six weeks. None of them required rewriting application code. None of them caused downtime. The hardest part wasn't the technology; it was the organizational change of making cost a visible, team-level metric.
If your EKS cluster is running on default Cluster Autoscaler settings with on-demand nodes and nobody knows how much each service costs, you're probably sitting on a similar 60-70% savings opportunity. Start with the resource audit. The numbers will do the convincing for you.
Top comments (0)