How does Kubernetes actually run an LLM on a GPU?
That was the focus of Day 18.
Kubernetes understands CPU and memory by default, but GPUs require additional components. We walked through the complete journey:
🔹 NFD (Node Feature Discovery) — Discovers GPU hardware on Kubernetes nodes.
🔹 GFD (GPU Feature Discovery) — Provides detailed NVIDIA GPU information such as model, memory, architecture, and MIG capability.
🔹 NVIDIA Device Plugin — Makes GPUs available as resources that Kubernetes workloads can request.
🔹 GPU Scheduling — We explored how nodeSelector, Node Affinity, and Taints/Tolerations help us choose the right GPU and protect expensive GPU nodes.
🔹 Dynamic Resource Allocation (DRA) — Instead of simply saying “Give me one GPU,” DRA allows workloads to describe what kind of device they need. We also looked at how DRA has evolved in Kubernetes 1.36.
🔹 NVIDIA GPU Operator — Automates much of the NVIDIA GPU software stack, including drivers, GPU discovery, the device plugin, container support, MIG management, and monitoring.
🔹 GPU Sharing — We compared Time Slicing and MIG and discussed when sharing an expensive GPU makes sense.
🔹 GPU Monitoring — We looked at how DCGM Exporter + Prometheus + Grafana can provide visibility into GPU utilization, memory, temperature, power, and health.
The big picture is:
Physical GPU → Discover → Expose → Schedule → Allocate → Run LLM → Monitor
For DevOps, SRE, and Platform Engineers moving into GenAI, GPUs are becoming another critical infrastructure resource we need to understand and manage.
Day 18/100 completed! 🚀
🚀 Want to learn GenAI from a DevOps Engineer's perspective?
🔗 https://ideaweaver.ai/#courses/genai-for-devops-engineers
📚 Follow Along Daily (Registration Required)
🎥 Day 18 (English): https://www.ideaweaver.ai/courses/100-days-of-genai-for-devops-english/lectures/66538878
🎥 Day 18(Hindi): https://www.ideaweaver.ai/courses/100-days-of-genai-for-devops-hindi/lectures/66538879
Top comments (0)