The average engineering team running Kubernetes is wasting between 30% and 50% of their cloud spend. Not because they're careless. Because every tutorial teaches you how to deploy to Kubernetes, and almost none of them teach you what the bill looks like six months later.
We've been running production Kubernetes clusters for client SaaS products for the past three years. Two of the systems we manage regularly come up in this conversation: a 250K daily active user EdTech platform, and a real-time GPS fleet tracking system with 30K active vehicle streams. We've made the mistakes. We've fixed them. Here's what actually moved the needle.
The #1 Source of Wasted Kubernetes Spend
Before you install any FinOps tool or switch to spot instances, look at your pod resource requests.
Resource requests are what Kubernetes uses to schedule your pods onto nodes. If your pods request more CPU or memory than they actually use, nodes fill up with "reserved but unused" capacity. You're paying for the node. You're not using the node.
# What we often find when we inherit a cluster
resources:
requests:
memory: "1Gi"
cpu: "500m"
limits:
memory: "2Gi"
cpu: "1000m"
That pod might actually be using 80MB of memory and 30m CPU at steady state. The node was reserved at 1Gi and 500m. You're paying for 12x the memory you need.
This is almost always the root cause, and it compounds fast. Ten pods like this means you're running a cluster that needs to be one-third the size it is.
How to find it:
# See actual CPU and memory use per pod
kubectl top pods --all-namespaces --sort-by=memory
# Compare to what's requested
kubectl describe pod <pod-name> | grep -A5 "Requests:"
Run that comparison. If your actual usage is consistently below 40% of your requests, you have a right-sizing problem.
Right-Sizing: The Vertical Pod Autoscaler Does the Math For You
Manually right-sizing 50 deployments is a week of work you don't have. The Vertical Pod Autoscaler (VPA) watches your actual usage over time and recommends better values.
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: api-server-vpa
namespace: production
spec:
targetRef:
apiVersion: "apps/v1"
kind: Deployment
name: api-server
updatePolicy:
updateMode: "Off" # Recommend only, don't auto-apply yet
resourcePolicy:
containerPolicies:
- containerName: "*"
minAllowed:
cpu: 10m
memory: 32Mi
maxAllowed:
cpu: 4
memory: 4Gi
Set updateMode: "Off" first. Let it observe for a week. Check the recommendations:
kubectl describe vpa api-server-vpa -n production
You'll see a recommendation block with lowerBound, target, and upperBound values. In our experience with production SaaS clusters, the recommended values are typically 50-70% lower than what teams initially set. Apply those recommendations to your deployment specs manually and you'll see a meaningful drop in node count within hours.
Spot Instances: The Biggest Lever You're Not Pulling
Spot instances (AWS) or preemptible VMs (GCP) can cut your compute cost by 60-90%. The tradeoff is that the cloud provider can terminate them with a 2-minute warning.
Most startups avoid spot because they're afraid of downtime. Here's the thing: if your application can't handle a pod being killed, you have a reliability problem that Kubernetes is supposed to solve for you anyway.
The right pattern: run your stateless workloads (API servers, workers, web apps) on spot, and your stateful workloads (databases, caches, queue processors that shouldn't be interrupted mid-job) on on-demand nodes.
# Node group for stateless workloads
apiVersion: eksctl.io/v1alpha5
kind: ClusterConfig
metadata:
name: production-cluster
region: us-east-1
nodeGroups:
- name: stateless-spot
instancesDistribution:
maxPrice: 0.08
instanceTypes: ["t3.medium", "t3a.medium", "t2.medium"]
onDemandBaseCapacity: 0
onDemandPercentageAboveBaseCapacity: 0 # 100% spot
spotAllocationStrategy: "capacity-optimized"
labels:
workload-type: stateless
taints:
- key: spot
value: "true"
effect: NoSchedule
Then in your deployment, tolerate spot nodes:
spec:
tolerations:
- key: spot
operator: Equal
value: "true"
effect: NoSchedule
nodeSelector:
workload-type: stateless
# Ensure pods spread across availability zones
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: api-server
Set a minimum of 2 replicas on anything running on spot. When a spot node is terminated, Kubernetes reschedules the pods in under 60 seconds. The short interruption is barely visible to users in practice.
For the GPS fleet tracking system, we run everything except the TimescaleDB cluster on spot. That single change (moving the Node.js API servers and workers to spot) cut the cluster compute cost significantly, and the system has been running cleanly through months of node terminations.
Namespace Quotas: Stopping Cost Sprawl Before It Starts
One engineering team's "quick test" deployment that nobody cleaned up is a classic story. Quotas at the namespace level stop it:
apiVersion: v1
kind: ResourceQuota
metadata:
name: staging-quota
namespace: staging
spec:
hard:
requests.cpu: "4"
requests.memory: 8Gi
limits.cpu: "8"
limits.memory: 16Gi
count/pods: "20"
count/persistentvolumeclaims: "5"
Now kubectl apply in staging fails gracefully when limits are hit instead of silently spawning more compute. Add LimitRanges so that pods without explicit resource declarations get sensible defaults:
apiVersion: v1
kind: LimitRange
metadata:
name: default-limits
namespace: staging
spec:
limits:
- default:
cpu: 200m
memory: 256Mi
defaultRequest:
cpu: 100m
memory: 128Mi
type: Container
This one change alone caught a situation for us where a misconfigured CI/CD pipeline was creating new pods on every commit without cleaning up the old ones. Within 20 minutes, the namespace quota was throwing errors. Without it, those pods would have run for weeks.
Horizontal Pod Autoscaling: Scale Down Aggressively at Night
If your SaaS product has predictable traffic patterns (and most do, with usage peaking in business hours and dropping at night), you should be scaling your pod count accordingly. HPA does this:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api-server-hpa
namespace: production
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api-server
minReplicas: 2
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 60
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 70
behavior:
scaleDown:
stabilizationWindowSeconds: 300 # Wait 5 min before scaling down
policies:
- type: Pods
value: 2
periodSeconds: 60
The scaleDown.stabilizationWindowSeconds is important. Without it, HPA can flap, scaling down too fast when traffic briefly dips, then scaling back up under the next request spike. A 5-minute stabilization window smooths this out.
For the EdTech platform, traffic drops to under 10% of peak between midnight and 6am. We set minReplicas: 2 for overnight and maxReplicas: 40 for peak. The cluster runs 4-6 nodes at night and 12-15 during the day. The arithmetic on that gap is significant over a month.
For workloads that should scale to near-zero (background workers, batch processors), look at KEDA. It can scale a deployment to 1 replica when a queue is empty, which HPA alone can't do.
Two Things Nobody Writes About
Image sizes compound infrastructure costs
A 2GB Docker image means a 2GB pull every time a pod starts on a new node. On Kubernetes at scale, that's bandwidth charges plus slower startup times that force you to keep more replicas "warm."
Multi-stage builds typically cut Node.js images from 800MB to under 150MB:
# Build stage
FROM node:20-alpine AS builder
WORKDIR /app
COPY package*.json ./
RUN npm ci --only=production
# Runtime stage
FROM node:20-alpine
WORKDIR /app
COPY --from=builder /app/node_modules ./node_modules
COPY --from=builder /app/dist ./dist
EXPOSE 3000
CMD ["node", "dist/index.js"]
That's an 80% reduction in image size. Faster pod starts, lower bandwidth, cheaper container registry storage. Small fix, visible in the bill.
Monitoring tools aren't free to run
Prometheus with default retention settings can quietly consume hundreds of gigabytes of persistent storage. Set retention policies:
# In your Prometheus values.yaml (Helm)
server:
retention: 15d # down from default 90d
retentionSize: "10GB"
persistentVolume:
size: 15Gi
You probably don't need 90 days of metrics in your operational monitoring store. Ship longer-term metrics to a cheaper storage layer (S3 + Thanos or Grafana Cloud's free tier) and keep hot storage lean.
Cluster Autoscaler: Make Sure It's Actually Scaling Down
Most teams enable Cluster Autoscaler (CA) for scale-up but don't configure the scale-down properly. The defaults are conservative: CA waits 10 minutes before removing a node, and only removes a node if it's been underutilized for that entire window. That's fine. But several settings override this and keep nodes alive longer than necessary.
# In your Cluster Autoscaler deployment args
- --scale-down-enabled=true
- --scale-down-delay-after-add=10m
- --scale-down-unneeded-time=10m
- --scale-down-utilization-threshold=0.5 # default is 0.5 (50%)
- --skip-nodes-with-local-storage=false # Set false if you're not running DaemonSets with storage
- --expander=least-waste # Don't always spin up the largest node
The expander=least-waste setting is the one people miss. By default, CA picks a node type based on random or a configured priority. least-waste picks the smallest node that fits the pending pod. For a startup with variable workload sizes, this can meaningfully reduce the average node size provisioned during scale-up events.
One pattern that trips people up: PodDisruptionBudgets that are set too conservatively block CA from evicting pods on underutilized nodes. If you have maxUnavailable: 0 on a 1-replica deployment, CA literally cannot drain that node. Check your PDBs:
kubectl get pdb --all-namespaces
Any PDB with ALLOWED DISRUPTIONS: 0 is a potential blocker for scale-down. Fix it to allow at least 1 disruption, or increase the replica count above 1.
The Quarterly Checklist
These are the specific questions we run through every quarter on the clusters we manage:
Resources:
- [ ] Any namespace where actual CPU/memory usage is below 40% of requests?
- [ ] Any deployments not touched in 60+ days that might be abandoned?
- [ ] Any PersistentVolumeClaims not attached to any pod?
Nodes:
- [ ] What percentage of workloads could run on spot without reliability risk?
- [ ] Are node sizes matched to workload shapes, or are we running
m5.largefor workloads that fit ont3.small?
Autoscaling:
- [ ] Do all production deployments have HPA configured?
- [ ] Do staging environments scale to near-zero outside business hours?
Storage:
- [ ] Are Prometheus and logging retention policies set?
- [ ] Any snapshots or backup volumes that can be cleaned up?
Run this once a quarter and you'll catch waste before it compounds into a surprise on your AWS bill.
Where to Start
If you're picking one thing to fix first: right-size your resource requests using VPA recommendations. It's the highest-leverage change and it costs you nothing except a week of observation data.
If you want the full picture (cost allocation across teams, charge-backs, proper FinOps practices), that's a bigger conversation. Our DevOps engineering team has run this process on production clusters serving millions of requests a day, and the patterns above are what we start with every time.
If you're at the stage of designing your infrastructure before it becomes expensive, our AWS cloud architecture services covers how we set up clusters to stay cost-effective from day one, including VPC layout, node group design, and GitOps pipelines.
The infrastructure decisions you make in the first six months of running Kubernetes are much easier to fix at the start than at $15,000/month. Most of the companies we work with wish they'd audited their resource requests in month two instead of month fourteen.
Yash Korat is the CEO and Co-Founder of Geminate Solutions, a custom software development company that has shipped 50+ products for startups across the US, UK, and Australia.
Originally published on Geminate Solutions.
Top comments (0)