I needed to give our internal platform tools access to Llama 3.3 70B and DeepSeek R1 without provisioning petabyte-scale PersistentVolumes for model weights. In this tutorial, we will build a lightweight LLM gateway, containerize it, and deploy it to a GPU-enabled Kubernetes cluster that calls Oxlo.ai for inference.
What you'll need
- A running Kubernetes cluster with GPU support (NVIDIA device plugin installed) and kubectl configured
- Docker
- Python 3.10 or newer
- The OpenAI SDK and FastAPI:
pip install openai fastapi uvicorn - An Oxlo.ai API key from https://portal.oxlo.ai
Step 1: Write the gateway service
I started by locking down the system prompt. I wanted the agent to act like a senior platform engineer, not a generic chatbot.
SYSTEM_PROMPT = """You are a senior platform engineer. Answer concisely.
Prefer code over prose. If you do not know something, say so."""
Then I wired it into a FastAPI app. I used the OpenAI SDK because Oxlo.ai exposes a fully compatible API.
import os
from fastapi import FastAPI
from pydantic import BaseModel
from openai import OpenAI
app = FastAPI()
client = OpenAI(
base_url="https://api.oxlo.ai/v1",
api_key=os.environ["OXLO_API_KEY"],
)
class ChatRequest(BaseModel):
message: str
@app.post("/chat")
async def chat(req: ChatRequest):
response = client.chat.completions.create(
model="llama-3.3-70b",
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": req.message},
],
)
return {"response": response.choices[0].message.content}
Step 2: Containerize the application
I kept the image small so it pulls quickly on new nodes. I used a slim Python base and installed the dependencies.
# requirements.txt
fastapi
uvicorn
openai
pydantic
# Dockerfile
FROM python:3.11-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY main.py .
EXPOSE 8000
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]
Step 3: Store the Oxlo.ai API key in Kubernetes
I created a Secret so the key never touches the container image.
apiVersion: v1
kind: Secret
metadata:
name: oxlo.ai-secret
namespace: default
type: Opaque
stringData:
OXLO_API_KEY: "YOUR_OXLO_API_KEY"
Step 4: Deploy to the GPU-enabled node pool
Our cluster runs GPU nodes for training and inference workloads. I added a nodeSelector and toleration so the gateway lands in that pool, but I did not request an NVIDIA device for the container itself. Oxlo.ai hosts the models, so we do not need to reserve local GPU memory for 70B weights.
apiVersion: apps/v1
kind: Deployment
metadata:
name: llm-gateway
namespace: default
spec:
replicas: 2
selector:
matchLabels:
app: llm-gateway
template:
metadata:
labels:
app: llm-gateway
spec:
nodeSelector:
accelerator: nvidia-gpu
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
containers:
- name: gateway
image: llm-gateway:latest
imagePullPolicy: IfNotPresent
ports:
- containerPort: 8000
env:
- name: OXLO_API_KEY
valueFrom:
secretKeyRef:
name: oxlo.ai-secret
key: OXLO_API_KEY
resources:
limits:
memory: "1Gi"
cpu: "1000m"
requests:
memory: "256Mi"
cpu: "100m"
Step 5: Expose the service
I added a ClusterIP service so other pods can reach the gateway at http://llm-gateway.default.svc.cluster.local.
apiVersion: v1
kind: Service
metadata:
name: llm-gateway
namespace: default
spec:
selector:
app: llm-gateway
ports:
- port: 80
targetPort: 8000
type: ClusterIP
Step 6: Apply the manifests
I built the image, pushed it to our internal registry, and applied the manifests.
docker build -t llm-gateway:latest .
# tag and push to your registry if necessary
kubectl apply -f secret.yaml
kubectl apply -f deployment.yaml
kubectl apply -f service.yaml
kubectl rollout status deployment/llm-gateway
Run it
I port-forwarded the service and sent a test request.
kubectl port-forward service/llm-gateway 8080:80
Then I called the endpoint with a question about Kubernetes probes.
curl -X POST http://localhost:8080/chat \
-H "Content-Type: application/json" \
-d '{"message": "Write a Python readiness probe for Kubernetes."}'
Oxlo.ai returned the following through our gateway:
{
"response": "from fastapi import FastAPI\nfrom fastapi.responses import PlainTextResponse\n\napp = FastAPI()\n\n@app.get('/healthz', response_class=PlainTextResponse)\ndef healthz():\n return 'ok'\n"
}
Wrap-up
You now have a stateless LLM gateway running on Kubernetes that calls Oxlo.ai. Two concrete next steps: add a HorizontalPodAutoscaler to scale replicas based on request rate, or swap the model to deepseek-v3.2 for coding tasks without changing any infrastructure.
Top comments (0)