DEV Community

Cover image for AI Workloads on Kubernetes 2026: The Production Stack (Kueue, KServe, vLLM, KubeRay)
saaro
saaro

Posted on Originally published at blog.saaro.net

AI Workloads on Kubernetes 2026: The Production Stack (Kueue, KServe, vLLM, KubeRay)

Two years ago, the question of whether Kubernetes is the right place for AI workloads was still open. By 2026, it has been answered: the production stack for AI/ML on Kubernetes has consolidated, the CNCF has graduated central projects, and the tooling landscape has matured. GPU clusters under Kubernetes are now part of the standard repertoire of platform engineering teams—whether for LLM inference, distributed training, or multi-tenant environments with multiple teams.

The 2026 Reference Architecture

A typical production stack for AI on Kubernetes consists of several layers that interlock seamlessly. The GPU infrastructure layer is formed by the NVIDIA GPU Operator (currently v26.3.3), which installs drivers, container runtime, device plugin, and DCGM monitoring with a single Helm chart and enables MIG partitioning as well as time-slicing. Above that lies the scheduling layer, where Kueue (CNCF, Kubernetes SIG project) handles admission control and quota management—crucial for multi-tenant clusters where multiple teams compete for GPUs. For distributed training with gang scheduling (where all pods of a job only start together), Volcano (CNCF Graduated) or the open-source KAI Scheduler is used. Often, Kueue and Volcano run as a layered system: Kueue holds back jobs when the quota is exhausted, while Volcano ensures atomic placement.

For inference, vLLM has established itself as the standard inference engine—with PagedAttention, continuous batching, and an OpenAI-compatible API. Around it, KServe (CNCF Incubating) provides a Kubernetes-native serving layer with autoscaling (including scale-to-zero), canary deployments, and traffic splitting. For models beyond the 70B parameter range (Llama 4 405B, DeepSeek V3), llm-d is increasingly used, a distributed inference framework with disaggregated serving that separates prefill and decode and distributes KV caches across nodes.

Training workloads typically run via Ray with the KubeRay operator, which provides RayCluster, RayJob, and RayService as native Kubernetes CRDs. KubeRay v1.7 (August 2026) brought a history server (beta), automatic mTLS certificate management, NetworkPolicy support, and Kubernetes RBAC-based authentication—a significant leap in production readiness.

GPU Scheduling: The Bottleneck No One Can Ignore

The biggest difference between classic microservices and AI workloads on Kubernetes is the handling of GPUs. The standard scheduler places pods individually—a distributed training job where seven out of eight worker pods start, but the eighth waits for a GPU, blocks the other seven GPUs at nearly zero utilization. This is exactly where gang scheduling (Volcano, KAI Scheduler) and quota queues (Kueue) come into play.Kueue boosts GPU utilization from 25–35% to 60–85% by enforcing fair-share queues across teams and enabling preemption for prioritized workloads. The combination of admission control and gang scheduling has now become the production standard for multi-tenant GPU clusters.

Topology-aware scheduling is also playing a growing role: pods placed on nodes with shared NVLink achieve three to five times lower inter-GPU latency than pods spanning separate network switches.

Multi-Tenancy and Cost Control

Shared GPU infrastructure is expensive—an unthrottled research workload can noticeably impact the latency of production inference endpoints. In practice, the standard has settled on Namespaces + RBAC + ResourceQuotas for logical partitioning, PriorityClasses for preemption, and Kueue ClusterQueues for cross-team fair-share policies.

On the cost side, Karpenter handles dynamic node provisioning: when a training job is pending, Karpenter provisions tailored spot or on-demand instances and scales them back down once the work is done. For GPU workloads that tolerate interruptions (training with checkpointing), spot instances are a massive cost lever—cloud providers advertise up to 90% discounts compared with on-demand prices.

Conclusion

The AI/ML stack on Kubernetes is production-ready and standardized in 2026. The core components—NVIDIA GPU Operator, Kueue, KServe, vLLM, Ray with KubeRay—are CNCF-graduated or on their way to it, and are supported by a large community. Platform teams that invest in this stack today benefit from portability across cloud providers and on-premises environments, fair GPU distribution among teams, and an ecosystem that is evolving rapidly—KubeRay v1.7, KServe's CNCF incubation, and the growing adoption of llm-d show that the direction is clear. Anyone who wants to run AI workloads seriously will no longer be able to avoid Kubernetes as the orchestration layer in 2026.

Sources

Top comments (0)