DEV Community

Geminate Solutions
Geminate Solutions

Posted on Originally published at geminatesolutions.com

Kubernetes Cost Optimization for Startups: What Actually Moves the Needle

The average engineering team running Kubernetes is wasting between 30% and 50% of their cloud spend. Not because they're careless. Because every tutorial teaches you how to deploy to Kubernetes, and almost none of them teach you what the bill looks like six months later.

We've been running production Kubernetes clusters for client SaaS products for the past three years. Two of the systems we manage regularly come up in this conversation: a 250K daily active user EdTech platform, and a real-time GPS fleet tracking system with 30K active vehicle streams. We've made the mistakes. We've fixed them. Here's what actually moved the needle.

The #1 Source of Wasted Kubernetes Spend

Before you install any FinOps tool or switch to spot instances, look at your pod resource requests.

Resource requests are what Kubernetes uses to schedule your pods onto nodes. If your pods request more CPU or memory than they actually use, nodes fill up with "reserved but unused" capacity. You're paying for the node. You're not using the node.

# What we often find when we inherit a cluster
resources:
  requests:
    memory: "1Gi"
    cpu: "500m"
  limits:
    memory: "2Gi"
    cpu: "1000m"
Enter fullscreen mode Exit fullscreen mode

That pod might actually be using 80MB of memory and 30m CPU at steady state. The node was reserved at 1Gi and 500m. You're paying for 12x the memory you need.

This is almost always the root cause, and it compounds fast. Ten pods like this means you're running a cluster that needs to be one-third the size it is.

How to find it:

# See actual CPU and memory use per pod
kubectl top pods --all-namespaces --sort-by=memory

# Compare to what's requested
kubectl describe pod <pod-name> | grep -A5 "Requests:"
Enter fullscreen mode Exit fullscreen mode

Run that comparison. If your actual usage is consistently below 40% of your requests, you have a right-sizing problem.

Right-Sizing: The Vertical Pod Autoscaler Does the Math For You

Manually right-sizing 50 deployments is a week of work you don't have. The Vertical Pod Autoscaler (VPA) watches your actual usage over time and recommends better values.

apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: api-server-vpa
  namespace: production
spec:
  targetRef:
    apiVersion: "apps/v1"
    kind: Deployment
    name: api-server
  updatePolicy:
    updateMode: "Off"  # Recommend only, don't auto-apply yet
  resourcePolicy:
    containerPolicies:
    - containerName: "*"
      minAllowed:
        cpu: 10m
        memory: 32Mi
      maxAllowed:
        cpu: 4
        memory: 4Gi
Enter fullscreen mode Exit fullscreen mode

Set updateMode: "Off" first. Let it observe for a week. Check the recommendations:

kubectl describe vpa api-server-vpa -n production
Enter fullscreen mode Exit fullscreen mode

You'll see a recommendation block with lowerBound, target, and upperBound values. In our experience with production SaaS clusters, the recommended values are typically 50-70% lower than what teams initially set. Apply those recommendations to your deployment specs manually and you'll see a meaningful drop in node count within hours.

Spot Instances: The Biggest Lever You're Not Pulling

Spot instances (AWS) or preemptible VMs (GCP) can cut your compute cost by 60-90%. The tradeoff is that the cloud provider can terminate them with a 2-minute warning.

Most startups avoid spot because they're afraid of downtime. Here's the thing: if your application can't handle a pod being killed, you have a reliability problem that Kubernetes is supposed to solve for you anyway.

The right pattern: run your stateless workloads (API servers, workers, web apps) on spot, and your stateful workloads (databases, caches, queue processors that shouldn't be interrupted mid-job) on on-demand nodes.

# Node group for stateless workloads
apiVersion: eksctl.io/v1alpha5
kind: ClusterConfig
metadata:
  name: production-cluster
  region: us-east-1
nodeGroups:
  - name: stateless-spot
    instancesDistribution:
      maxPrice: 0.08
      instanceTypes: ["t3.medium", "t3a.medium", "t2.medium"]
      onDemandBaseCapacity: 0
      onDemandPercentageAboveBaseCapacity: 0  # 100% spot
      spotAllocationStrategy: "capacity-optimized"
    labels:
      workload-type: stateless
    taints:
      - key: spot
        value: "true"
        effect: NoSchedule
Enter fullscreen mode Exit fullscreen mode

Then in your deployment, tolerate spot nodes:

spec:
  tolerations:
  - key: spot
    operator: Equal
    value: "true"
    effect: NoSchedule
  nodeSelector:
    workload-type: stateless
  # Ensure pods spread across availability zones
  topologySpreadConstraints:
  - maxSkew: 1
    topologyKey: topology.kubernetes.io/zone
    whenUnsatisfiable: DoNotSchedule
    labelSelector:
      matchLabels:
        app: api-server
Enter fullscreen mode Exit fullscreen mode

Set a minimum of 2 replicas on anything running on spot. When a spot node is terminated, Kubernetes reschedules the pods in under 60 seconds. The short interruption is barely visible to users in practice.

For the GPS fleet tracking system, we run everything except the TimescaleDB cluster on spot. That single change (moving the Node.js API servers and workers to spot) cut the cluster compute cost significantly, and the system has been running cleanly through months of node terminations.

Namespace Quotas: Stopping Cost Sprawl Before It Starts

One engineering team's "quick test" deployment that nobody cleaned up is a classic story. Quotas at the namespace level stop it:

apiVersion: v1
kind: ResourceQuota
metadata:
  name: staging-quota
  namespace: staging
spec:
  hard:
    requests.cpu: "4"
    requests.memory: 8Gi
    limits.cpu: "8"
    limits.memory: 16Gi
    count/pods: "20"
    count/persistentvolumeclaims: "5"
Enter fullscreen mode Exit fullscreen mode

Now kubectl apply in staging fails gracefully when limits are hit instead of silently spawning more compute. Add LimitRanges so that pods without explicit resource declarations get sensible defaults:

apiVersion: v1
kind: LimitRange
metadata:
  name: default-limits
  namespace: staging
spec:
  limits:
  - default:
    cpu: 200m
    memory: 256Mi
  defaultRequest:
    cpu: 100m
    memory: 128Mi
  type: Container
Enter fullscreen mode Exit fullscreen mode

This one change alone caught a situation for us where a misconfigured CI/CD pipeline was creating new pods on every commit without cleaning up the old ones. Within 20 minutes, the namespace quota was throwing errors. Without it, those pods would have run for weeks.

Horizontal Pod Autoscaling: Scale Down Aggressively at Night

If your SaaS product has predictable traffic patterns (and most do, with usage peaking in business hours and dropping at night), you should be scaling your pod count accordingly. HPA does this:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: api-server-hpa
  namespace: production
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: api-server
  minReplicas: 2
  maxReplicas: 20
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 60
  - type: Resource
    resource:
      name: memory
      target:
        type: Utilization
        averageUtilization: 70
  behavior:
    scaleDown:
      stabilizationWindowSeconds: 300  # Wait 5 min before scaling down
      policies:
      - type: Pods
        value: 2
        periodSeconds: 60
Enter fullscreen mode Exit fullscreen mode

The scaleDown.stabilizationWindowSeconds is important. Without it, HPA can flap, scaling down too fast when traffic briefly dips, then scaling back up under the next request spike. A 5-minute stabilization window smooths this out.

For the EdTech platform, traffic drops to under 10% of peak between midnight and 6am. We set minReplicas: 2 for overnight and maxReplicas: 40 for peak. The cluster runs 4-6 nodes at night and 12-15 during the day. The arithmetic on that gap is significant over a month.

For workloads that should scale to near-zero (background workers, batch processors), look at KEDA. It can scale a deployment to 1 replica when a queue is empty, which HPA alone can't do.

Two Things Nobody Writes About

Image sizes compound infrastructure costs

A 2GB Docker image means a 2GB pull every time a pod starts on a new node. On Kubernetes at scale, that's bandwidth charges plus slower startup times that force you to keep more replicas "warm."

Multi-stage builds typically cut Node.js images from 800MB to under 150MB:

# Build stage
FROM node:20-alpine AS builder
WORKDIR /app
COPY package*.json ./
RUN npm ci --only=production

# Runtime stage
FROM node:20-alpine
WORKDIR /app
COPY --from=builder /app/node_modules ./node_modules
COPY --from=builder /app/dist ./dist
EXPOSE 3000
CMD ["node", "dist/index.js"]
Enter fullscreen mode Exit fullscreen mode

That's an 80% reduction in image size. Faster pod starts, lower bandwidth, cheaper container registry storage. Small fix, visible in the bill.

Monitoring tools aren't free to run

Prometheus with default retention settings can quietly consume hundreds of gigabytes of persistent storage. Set retention policies:

# In your Prometheus values.yaml (Helm)
server:
  retention: 15d  # down from default 90d
  retentionSize: "10GB"
  persistentVolume:
    size: 15Gi
Enter fullscreen mode Exit fullscreen mode

You probably don't need 90 days of metrics in your operational monitoring store. Ship longer-term metrics to a cheaper storage layer (S3 + Thanos or Grafana Cloud's free tier) and keep hot storage lean.

Cluster Autoscaler: Make Sure It's Actually Scaling Down

Most teams enable Cluster Autoscaler (CA) for scale-up but don't configure the scale-down properly. The defaults are conservative: CA waits 10 minutes before removing a node, and only removes a node if it's been underutilized for that entire window. That's fine. But several settings override this and keep nodes alive longer than necessary.

# In your Cluster Autoscaler deployment args
- --scale-down-enabled=true
- --scale-down-delay-after-add=10m
- --scale-down-unneeded-time=10m
- --scale-down-utilization-threshold=0.5  # default is 0.5 (50%)
- --skip-nodes-with-local-storage=false   # Set false if you're not running DaemonSets with storage
- --expander=least-waste                  # Don't always spin up the largest node
Enter fullscreen mode Exit fullscreen mode

The expander=least-waste setting is the one people miss. By default, CA picks a node type based on random or a configured priority. least-waste picks the smallest node that fits the pending pod. For a startup with variable workload sizes, this can meaningfully reduce the average node size provisioned during scale-up events.

One pattern that trips people up: PodDisruptionBudgets that are set too conservatively block CA from evicting pods on underutilized nodes. If you have maxUnavailable: 0 on a 1-replica deployment, CA literally cannot drain that node. Check your PDBs:

kubectl get pdb --all-namespaces
Enter fullscreen mode Exit fullscreen mode

Any PDB with ALLOWED DISRUPTIONS: 0 is a potential blocker for scale-down. Fix it to allow at least 1 disruption, or increase the replica count above 1.

The Quarterly Checklist

These are the specific questions we run through every quarter on the clusters we manage:

Resources:

  • [ ] Any namespace where actual CPU/memory usage is below 40% of requests?
  • [ ] Any deployments not touched in 60+ days that might be abandoned?
  • [ ] Any PersistentVolumeClaims not attached to any pod?

Nodes:

  • [ ] What percentage of workloads could run on spot without reliability risk?
  • [ ] Are node sizes matched to workload shapes, or are we running m5.large for workloads that fit on t3.small?

Autoscaling:

  • [ ] Do all production deployments have HPA configured?
  • [ ] Do staging environments scale to near-zero outside business hours?

Storage:

  • [ ] Are Prometheus and logging retention policies set?
  • [ ] Any snapshots or backup volumes that can be cleaned up?

Run this once a quarter and you'll catch waste before it compounds into a surprise on your AWS bill.

Where to Start

If you're picking one thing to fix first: right-size your resource requests using VPA recommendations. It's the highest-leverage change and it costs you nothing except a week of observation data.

If you want the full picture (cost allocation across teams, charge-backs, proper FinOps practices), that's a bigger conversation. Our DevOps engineering team has run this process on production clusters serving millions of requests a day, and the patterns above are what we start with every time.

If you're at the stage of designing your infrastructure before it becomes expensive, our AWS cloud architecture services covers how we set up clusters to stay cost-effective from day one, including VPC layout, node group design, and GitOps pipelines.

The infrastructure decisions you make in the first six months of running Kubernetes are much easier to fix at the start than at $15,000/month. Most of the companies we work with wish they'd audited their resource requests in month two instead of month fourteen.


Yash Korat is the CEO and Co-Founder of Geminate Solutions, a custom software development company that has shipped 50+ products for startups across the US, UK, and Australia.

Originally published on Geminate Solutions.

Top comments (0)