A field survey of the Kubernetes + LLM inference stack in 2026 — the problems that now have serious players, and the edges that are still genuinely underserved.
If you run large language models on Kubernetes, you've probably noticed the ground shifting fast. A year ago most of the hard problems were open. In 2026 a lot of them have serious, well-funded players — which changes the calculus if you're deciding where to build. Competing with a GA project is very different from pioneering an empty space.
This is a map of the territory: what's crowded, and where I think the real gaps still are. I'm deliberately not pitching any of these as "my novel architecture" — several have academic papers or shipping projects attached. The point is to show where the open problems are so you can decide where your effort pays off.
What's already crowded
These are the areas where the obvious pain points now have mature answers. Building here means competing, not pioneering.
KV-cache-aware routing / model-aware gateways. Largely solved. The Gateway API Inference Extension (a kubernetes-sigs project) defines a standard InferencePool routing contract, and llm-d, NVIDIA Dynamo, and GKE Inference Gateway (which went GA in 2026) all do prefix-cache affinity and cache-aware routing on top of it.
Prefill/decode disaggregation. NVIDIA Dynamo and llm-d cover this well.
Fractional GPU sharing (MIG / MPS / time-slicing). Run:ai, cast.ai, the NVIDIA GPU Operator, and vCluster all address it.
GPU scheduling primitives (DRA, gang scheduling). KAI Scheduler, Grove, Kueue, and Volcano are all in play, with Dynamic Resource Allocation graduating in core Kubernetes. Crowded, though still messy at the edges.
If your idea lives in one of those buckets, go in with eyes open: you're entering a race.
Where the real gaps still are
Three edges still look genuinely underserved, and each is shaped like a Kubernetes-native problem — a CRD plus a controller plus some node-level machinery.
1. Model-weight distribution as a first-class Kubernetes primitive
This is the one I find most interesting. Model weights are often 100+ GB, and today they're still delivered largely by ad-hoc downloads from object storage — without the pull-caching, digest addressing, and content verification that Kubernetes already gives you for container images. That gap is a major driver of the two-to-ten-minute cold starts that make scale-to-zero economically painful for inference, since every new replica pays that cost before serving its first request.
The literature backs up the problem. An empirical study (arXiv 2607.16596, "Cold-Start Model Delivery in Kubernetes Inference Serving") measures the delivery paths — modelcar sidecars, native image volumes, object-storage download — on artifacts sized to 1B-, 7B-, and 70B-class weights, and a companion paper (arXiv 2609.20874) frames the same two-to-ten-minute startup as the reason reactive autoscaling is structurally late. So the pain is validated; what's missing is a first-class primitive for it. Note that these papers study and characterize the problem — they don't ship the CRD-shaped solution below, and neither do I claim to have invented it.
The interesting shape of a solution: a ModelArtifact / ModelCache CRD, a controller, and a node-level DaemonSet that treats weights like OCI-addressable, content-verified, node-locally-cached blobs — with topology-aware peer-to-peer prewarming (pull from a neighbor node's cache over the fast fabric instead of re-pulling from object storage), plus scheduler hints so pods land where the weights already live. NVIDIA Dynamo's ModelExpress already does some of this (one worker publishes weights, others pull over NIXL/RDMA), so even here you'd be refining a direction others have started, not working in a vacuum.
Why it's a good build target: it's infrastructure-shaped, CRD-heavy, and not yet a solved commodity. You can prototype the caching, P2P, and scheduler-hint logic and measure cold-start reduction on modest hardware — or even with simulated weights — without a large GPU budget.
2. SLO-driven inference autoscaling that understands request heterogeneity
A caution up front: "SLO-aware inference autoscaling" as a headline is already a shipped commodity — AWS, KEDA, and academia all have versions of it. So this is only interesting if you go narrower than the headline.
Current autoscaling scales on queue depth or KV-cache utilization. But the Cascade paper (arXiv 2608.06557, "Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving") shows that requests under the same SLO can differ in urgency by a wide margin — a short chat versus a long agentic reasoning trace — and that you can direct queueing and data-movement overhead toward the requests that can absorb it. Related systems like Scorpio (arXiv 2505.23022) exploit the same SLO heterogeneity for admission and scheduling, so this is an active research line, not an empty one. The defensible wedge is per-request-class latency budgeting plus admission control: an InferenceSLO CRD where you declare per-class latency budgets, and a controller that combines a batch-latency predictor with HPA/KEDA to make admission and scaling decisions on predicted goodput rather than raw queue length.
This one is more research-flavored and genuinely publishable — but only if you build and benchmark it against the existing SLO-aware baselines. Without numbers, it reads as a restatement of what's already shipping.
3. Multi-LoRA adapter lifecycle as a CRD
Serving hundreds of tenant-specific LoRA adapters over one base model is now a real production pattern (arXiv 2511.22880), and vLLM and Dynamo support runtime adapter loading. What's still hand-rolled is lifecycle management: which adapters are hot, cache eviction by usage, routing requests to workers that already hold the adapter, and VRAM budgeting. Dynamo has some declarative Kubernetes support here, but it's early.
The shape: a LoRAAdapter + AdapterPool CRD that manages adapter placement, usage-based eviction, and adapter-affinity routing hints — essentially "KV-cache-aware routing," but for adapters.
How I'd choose
If you want a portfolio or OSS operator project: idea #1. It's clearly the most infrastructure-shaped, there's a fresh paper validating the problem, and it's demonstrable without a big GPU budget.
If you'd rather write something paper-worthy: idea #2, with the per-request-class framing and real benchmarks.
Idea #3 is a reasonable middle ground.
A closing note on honesty, because it matters in this space: none of these are empty green fields. Each sits next to active prior art, and dev.to readers who work on inference will know it. The value here isn't a claim of novelty — it's a clear-eyed map of where the race is already crowded and where there's still room to build something useful.
Sources and prior art
Content was rephrased from the sources below for licensing compliance. Verify each link before publishing.
- Gateway API Inference Extension — kubernetes-sigs/gateway-api-inference-extension (github.com)
- GKE Inference Gateway — GA, Google Cloud blog, 2026 (cloud.google.com)
- llm-d
- NVIDIA Dynamo; ModelExpress weight distribution (docs.nvidia.com)
- Run:ai, cast.ai, NVIDIA GPU Operator, vCluster (GPU sharing)
- KAI Scheduler, Grove, Kueue, Volcano; Dynamic Resource Allocation in Kubernetes
- arXiv 2607.16596 — "Cold-Start Model Delivery in Kubernetes Inference Serving: An Empirical Study of OCI-Based Distribution and Its Integrity" (G. Kliukovkin)
- arXiv 2609.20874 — "Decomposing Predictive Kubernetes Autoscaling for LLM Serving Under Long Startup Delays"
- arXiv 2608.06557 — "Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving"
- arXiv 2505.23022 — "Scorpio: Serving Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference"
- arXiv 2511.22880 — "Serving Heterogeneous LoRA Adapters in Distributed LLM Inference Systems"
Top comments (0)