I intentionally over-provisioned a Kubernetes application just to see whether CAST AI could detect it, recommend better resource requests, and actually optimize it.
Spoiler: it did—but the biggest lesson wasn't about cost savings. It was about understanding the difference between resource requests, actual usage, and infrastructure cost.
🤔 Why I Built This Lab
When we deploy applications to Kubernetes, we usually write something like this:
resources:
requests:
cpu: "1000m"
memory: "1Gi"
limits:
cpu: "1500m"
memory: "1500Mi"
But here's the question:
How do you know if those numbers are correct?
Most of us guess.
Sometimes we copy them from another project.
Sometimes we simply over-provision because we don't want the application to crash.
That's exactly what I did for this experiment.
🏗️ The Lab Setup
I built a very small application instead of using a complex microservice.
Tech Stack
FastAPI
Docker
Google Kubernetes Engine (GKE)
Google Artifact Registry
CAST AI
The deployment flow looked like this:
FastAPI
│
Docker Image
│
Artifact Registry
│
GKE Cluster
│
Deployment (2 Pods)
│
ClusterIP Service
│
CAST AI
☕ The Coffee API
The application only exposes two endpoints.
@app.get("/")
def home():
return {
"message": "Coffee Shop API",
"hostname": socket.gethostname()
}
@app.get("/coffee")
def coffee():
return {
"coffee": "Cappuccino",
"price": 120
}
Nothing fancy.
The goal wasn't the application.
The goal was the infrastructure.
🎯 Intentionally Over-Provisioning the Pods
Each Pod requested:
resources:
requests:
cpu: "1000m"
memory: "1Gi"
and I deployed 2 replicas.
So Kubernetes reserved:
2 CPU
2 GiB Memory
for an application that barely did anything.
🚀 Deploying to GKE
After building the Docker image, I pushed it to Artifact Registry and deployed it to GKE.
Everything looked healthy.
2 Pods
Running
Ready
At this point, Kubernetes had no idea the application was oversized.
It simply respected the resource requests I had defined.
🔍 Connecting CAST AI
I connected the cluster to CAST AI.
Interestingly, onboarding wasn't completely smooth.
I ran into several issues:
missing gke-gcloud-auth-plugin
expired CAST AI authentication
missing IAM permissions
Kubernetes RBAC
After fixing those, CAST AI started collecting workload metrics.
📊 The Interesting Part
After some time, CAST AI analyzed my workload.
My deployment requested:
CPU
1 CPU per Pod
Memory
1 GiB per Pod
But actual usage looked more like:
CPU
≈ 3-4m
Memory
≈ 30 MiB
That was a huge difference.
💡 CAST AI Recommendation
Instead of recommending something extreme like 4m CPU, CAST AI suggested:
Resource Before Recommendation
CPU 1 CPU 750m
Memory 1 GiB 768 MiB
It also estimated around 25% workload cost reduction.
Screenshot
(Insert your CAST AI optimization screenshot here)
🤔 Why Not Just Recommend 4m CPU?
This was actually my favorite learning.
If the application is only using 4m CPU...
...why recommend 750m?
Because optimization isn't simply:
request = current usage
Production workloads experience:
startup spikes
traffic bursts
background jobs
unpredictable load
A good optimizer leaves enough headroom for those situations.
✅ Verifying It From Kubernetes
I didn't want to trust only the dashboard.
So I verified everything directly from Kubernetes.
kubectl get pod \
-o jsonpath='{.status.containerStatuses[0].allocatedResources}'
Result:
{
"cpu":"750m",
"memory":"768Mi"
}
That was pretty satisfying.
🧠 Biggest Takeaways
This experiment completely changed how I think about Kubernetes resources.
I learned that these are not the same thing:
Requested Resources
↓
Actual Usage
↓
Recommended Resources
↓
Allocated Resources
↓
Infrastructure Cost
Those five numbers often get mixed together.
Understanding the difference is the first step toward Kubernetes cost optimization.
Final Thoughts
Before this project I thought Kubernetes cost optimization mostly meant:
"Find cheaper virtual machines."
Now I'd describe it differently:
Start by making sure your workloads request only the resources they actually need.
Node optimization comes later.
Workload optimization comes first.
📚 Project Repository
I documented the complete lab including:
GKE setup
CAST AI onboarding
Docker deployment
Kubernetes networking
Cost monitoring
Workload optimization
Troubleshooting
👉 GitHub:
https://github.com/Upshivam786/castai-gke-workload-optimization-lab
🙌 Thanks for Reading
If you've experimented with Kubernetes cost optimization tools like CAST AI, Karpenter, Goldilocks, or Vertical Pod Autoscaler, I'd love to hear about your experience.
Happy learning! 🚀
Tags:
Top comments (0)