DEV Community

shashank ms
shashank ms

Posted on

Deploying LLM Models on Kubernetes Clusters with GPU Support

I needed to give our internal platform tools access to Llama 3.3 70B and DeepSeek R1 without provisioning petabyte-scale PersistentVolumes for model weights. In this tutorial, we will build a lightweight LLM gateway, containerize it, and deploy it to a GPU-enabled Kubernetes cluster that calls Oxlo.ai for inference.

What you'll need

  • A running Kubernetes cluster with GPU support (NVIDIA device plugin installed) and kubectl configured
  • Docker
  • Python 3.10 or newer
  • The OpenAI SDK and FastAPI: pip install openai fastapi uvicorn
  • An Oxlo.ai API key from https://portal.oxlo.ai

Step 1: Write the gateway service

I started by locking down the system prompt. I wanted the agent to act like a senior platform engineer, not a generic chatbot.

SYSTEM_PROMPT = """You are a senior platform engineer. Answer concisely.
Prefer code over prose. If you do not know something, say so."""

Then I wired it into a FastAPI app. I used the OpenAI SDK because Oxlo.ai exposes a fully compatible API.

import os
from fastapi import FastAPI
from pydantic import BaseModel
from openai import OpenAI

app = FastAPI()

client = OpenAI(
    base_url="https://api.oxlo.ai/v1",
    api_key=os.environ["OXLO_API_KEY"],
)

class ChatRequest(BaseModel):
    message: str

@app.post("/chat")
async def chat(req: ChatRequest):
    response = client.chat.completions.create(
        model="llama-3.3-70b",
        messages=[
            {"role": "system", "content": SYSTEM_PROMPT},
            {"role": "user", "content": req.message},
        ],
    )
    return {"response": response.choices[0].message.content}

Step 2: Containerize the application

I kept the image small so it pulls quickly on new nodes. I used a slim Python base and installed the dependencies.

# requirements.txt
fastapi
uvicorn
openai
pydantic
# Dockerfile
FROM python:3.11-slim

WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

COPY main.py .
EXPOSE 8000

CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]

Step 3: Store the Oxlo.ai API key in Kubernetes

I created a Secret so the key never touches the container image.

apiVersion: v1
kind: Secret
metadata:
  name: oxlo.ai-secret
  namespace: default
type: Opaque
stringData:
  OXLO_API_KEY: "YOUR_OXLO_API_KEY"

Step 4: Deploy to the GPU-enabled node pool

Our cluster runs GPU nodes for training and inference workloads. I added a nodeSelector and toleration so the gateway lands in that pool, but I did not request an NVIDIA device for the container itself. Oxlo.ai hosts the models, so we do not need to reserve local GPU memory for 70B weights.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: llm-gateway
  namespace: default
spec:
  replicas: 2
  selector:
    matchLabels:
      app: llm-gateway
  template:
    metadata:
      labels:
        app: llm-gateway
    spec:
      nodeSelector:
        accelerator: nvidia-gpu
      tolerations:
        - key: nvidia.com/gpu
          operator: Exists
          effect: NoSchedule
      containers:
        - name: gateway
          image: llm-gateway:latest
          imagePullPolicy: IfNotPresent
          ports:
            - containerPort: 8000
          env:
            - name: OXLO_API_KEY
              valueFrom:
                secretKeyRef:
                  name: oxlo.ai-secret
                  key: OXLO_API_KEY
          resources:
            limits:
              memory: "1Gi"
              cpu: "1000m"
            requests:
              memory: "256Mi"
              cpu: "100m"

Step 5: Expose the service

I added a ClusterIP service so other pods can reach the gateway at http://llm-gateway.default.svc.cluster.local.

apiVersion: v1
kind: Service
metadata:
  name: llm-gateway
  namespace: default
spec:
  selector:
    app: llm-gateway
  ports:
    - port: 80
      targetPort: 8000
  type: ClusterIP

Step 6: Apply the manifests

I built the image, pushed it to our internal registry, and applied the manifests.

docker build -t llm-gateway:latest .
# tag and push to your registry if necessary
kubectl apply -f secret.yaml
kubectl apply -f deployment.yaml
kubectl apply -f service.yaml
kubectl rollout status deployment/llm-gateway

Run it

I port-forwarded the service and sent a test request.

kubectl port-forward service/llm-gateway 8080:80

Then I called the endpoint with a question about Kubernetes probes.

curl -X POST http://localhost:8080/chat \
  -H "Content-Type: application/json" \
  -d '{"message": "Write a Python readiness probe for Kubernetes."}'

Oxlo.ai returned the following through our gateway:

{
  "response": "from fastapi import FastAPI\nfrom fastapi.responses import PlainTextResponse\n\napp = FastAPI()\n\n@app.get('/healthz', response_class=PlainTextResponse)\ndef healthz():\n    return 'ok'\n"
}

Wrap-up

You now have a stateless LLM gateway running on Kubernetes that calls Oxlo.ai. Two concrete next steps: add a HorizontalPodAutoscaler to scale replicas based on request rate, or swap the model to deepseek-v3.2 for coding tasks without changing any infrastructure.

Top comments (0)