DEV Community

Cover image for Deploying an open-source AI agent platform to Kubernetes: the honest one-command version
Anis Meziani
Anis Meziani

Posted on

Deploying an open-source AI agent platform to Kubernetes: the honest one-command version

Every "deploy X to Kubernetes in one command" article is technically true and practically a lie. The helm install really is one line. What the post skips is the empty cluster with no storage class, no ingress controller, no cert-manager, and no DNS. That 90% is what stands between you and a working URL. So here's the honest version for ApowerB, the open-source agent platform: the chart genuinely installs in one command, and here is everything around it, traps included.


What the chart actually deploys

ApowerB ships an official Helm chart (apowerb-chart, OCI, cosign-signed) that stands up the whole stack, not just the app:

  • Backend and Next.js frontend: the agent runtime and its UI
  • PostgreSQL: where your agents live (they're rows, materialized into executable modules at load time)
  • th2etl: pipeline and orchestration, with a seed job
  • th2pulse: log and trace storage
  • otel-collector: telemetry routing
  • th2forecast: optional forecasting engine (disabled by default)
  • A PersistentVolumeClaim for uploads, RAG artifacts and agent pools

That's a platform, not a container. Which is exactly why the cluster prerequisites below matter more than for a toy app.


The fast path (local or dev cluster)

If you already have a cluster with a working storage class, two commands get you running.

1. Generate the secrets. The chart refuses to start without them, on purpose:

umask 077
cat > values-secrets.yaml <<YAML
backend:
  env:
    encryptKey: "$(openssl rand -base64 32)"
th2etl:
  apiKey: "$(openssl rand -hex 32)"
th2pulse:
  ingestToken: "$(openssl rand -hex 32)"
  queryToken: "$(openssl rand -hex 32)"
postgres:
  password: "$(openssl rand -hex 16)"
YAML
chmod 600 values-secrets.yaml
Enter fullscreen mode Exit fullscreen mode

2. Install:

helm upgrade --install apowerb \
  oci://registry-1.docker.io/apowerb/apowerb-chart --version 0.4.23 \
  --namespace apowerb --create-namespace \
  --values values-secrets.yaml --timeout 10m
Enter fullscreen mode Exit fullscreen mode

Then port-forward and open the UI:

kubectl -n apowerb port-forward svc/apowerb-frontend 13000:3000 &
# http://localhost:13000
Enter fullscreen mode Exit fullscreen mode

Add a model key in the UI and you have a working agent platform on your cluster. That's the "one command" part, and it's real.

The one rule that bites everyone: pass --values values-secrets.yaml on every helm upgrade. Drop it and you rotate your own credentials out from under a running database. Keep that file safe and version it in your secret store, not in Git.


The honest production path (empty cluster to HTTPS URL)

A real cluster rarely has everything wired. Here's the full walk, which is where most tutorials go quiet.

Prerequisites

Kubernetes 1.28+. Three nodes at 4 vCPU / 8 GB give comfortable headroom. Check the ground first:

kubectl get nodes          # all Ready
kubectl get storageclass   # must not be empty
kubectl get ingressclass   # must exist
Enter fullscreen mode Exit fullscreen mode

Step 1: Storage that actually binds

Lean managed clusters often ship with no default storage class. If get storageclass is empty, install one:

curl -sL -o local-path-storage.yaml \
  https://raw.githubusercontent.com/rancher/local-path-provisioner/v0.0.37/deploy/local-path-storage.yaml
kubectl apply -f local-path-storage.yaml
kubectl -n local-path-storage rollout status deploy/local-path-provisioner --timeout=120s

kubectl patch storageclass local-path \
  -p '{"metadata":{"annotations":{"storageclass.kubernetes.io/is-default-class":"true"}}}'
Enter fullscreen mode Exit fullscreen mode

Trap: most storage classes are WaitForFirstConsumer, so a lone PVC stays Pending and looks broken when it's fine. Test with a PVC and a pod together:

kubectl apply -f - <<'YAML'
apiVersion: v1
kind: PersistentVolumeClaim
metadata: {name: probe-pvc}
spec:
  accessModes: ["ReadWriteOnce"]
  resources: {requests: {storage: 1Gi}}
---
apiVersion: v1
kind: Pod
metadata: {name: probe-pod}
spec:
  restartPolicy: Never
  containers:
    - name: writer
      image: busybox:1.37
      command: ["sh","-c","echo ok > /data/marker && cat /data/marker && sleep 20"]
      volumeMounts: [{name: data, mountPath: /data}]
  volumes:
    - {name: data, persistentVolumeClaim: {claimName: probe-pvc}}
YAML
sleep 25 && kubectl get pvc probe-pvc && kubectl logs probe-pod
kubectl delete pvc probe-pvc pod probe-pod
Enter fullscreen mode Exit fullscreen mode

Want Bound and ok in the logs. If this fails, stop. Nothing else will work.

Step 2: Ingress

kubectl get ingressclass
kubectl get svc -A | grep LoadBalancer
Enter fullscreen mode Exit fullscreen mode

Note the class name (Traefik and ingress-nginx are the usual managed add-ons) and the EXTERNAL-IP for DNS.

Step 3: cert-manager and a staging issuer

Let's Encrypt rate-limits at 5 failures/hour, so always start on staging:

apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt-staging
spec:
  acme:
    server: https://acme-staging-v02.api.letsencrypt.org/directory
    email: you@example.com
    privateKeySecretRef:
      name: letsencrypt-staging-account-key
    solvers:
      - http01:
          ingress:
            ingressClassName: traefik   # your actual class from Step 2
Enter fullscreen mode Exit fullscreen mode

Trap that costs an afternoon: use ingressClassName, not the deprecated class annotation. Otherwise the ACME challenge Ingress gets no recognised class and the certificate sits at Ready: False forever.

Step 4: Deploy with ingress and TLS

Same install command, now with public access turned on:

helm upgrade --install apowerb \
  oci://registry-1.docker.io/apowerb/apowerb-chart --version 0.4.23 \
  --namespace apowerb --values values-secrets.yaml \
  --set ingress.enabled=true \
  --set ingress.className=traefik \
  --set ingress.host=apowerb.example.com \
  --set ingress.tlsEnabled=true \
  --set ingress.annotations."cert-manager\.io/cluster-issuer"=letsencrypt-staging
Enter fullscreen mode Exit fullscreen mode

Create the first admin while you're at it:

  --set superadmin.email=admin@example.com \
  --set superadmin.password='<a strong one>'
Enter fullscreen mode Exit fullscreen mode

Step 5: DNS and verify

Point an A record at the EXTERNAL-IP, then:

kubectl -n apowerb get certificate
kubectl -n apowerb describe certificate apowerb-tls
Enter fullscreen mode Exit fullscreen mode

Stuck at Ready: False? It's almost always one of: DNS not resolving to the load balancer, port 80 closed (HTTP-01 needs it), the class vs ingressClassName trap, or a staging rate limit. Once it's True, swap the issuer to production (https://acme-v02.api.letsencrypt.org/directory) and reapply.


Configuring the platform

The same values mechanism covers the rest:

  • Default LLM: defaultLlm.model and defaultLlm.apiKey for agents without their own key (LiteLLM under the hood, so OpenAI / Anthropic / Mistral / Gemini / local all work)
  • Integrations: GOOGLE_INTEGRATION_* (Drive, Gmail, Calendar), MICROSOFT_INTEGRATION_* (Outlook, webhooks, mail), SMTP_*
  • Storage: storage.mode and S3_* to push bi_store, uploads, artifacts_store and agents_pool to object storage instead of the PVC

Verdict

The install itself is genuinely one command against an OCI chart, and the chart is honest: it won't boot without real secrets, and it bundles the orchestration, log store and telemetry rather than leaving them as "exercises for the reader." The work that remains isn't ApowerB's, it's the cluster's: storage, ingress, cert-manager, DNS. And that's true of anything serious you deploy.

If you want to see the easy version first, the Docker Compose stack comes up locally in a few commands. But if you're putting agents in front of real users, the Helm chart is the path, and now you've seen all of it, traps included.


Links

Deployed it on something interesting? Drop your cluster setup in the comments.

Top comments (2)

Collapse
 
adil_harcha_3942e1aa958eb profile image
Adil Harcha •

Thanks for this refreshing and refreshingly honest write-up! 👏 Finally, an article that breaks down the myth of the magical "one-command deploy" without glossing over the actual cluster jungle underneath.

Your warning about passing --values values-secrets.yaml on every single upgrade is gold: that's exactly the kind of silent trap that wreaks havoc in production if it isn't automated away via an operator or a tool like External Secrets / Vault.

A couple of quick questions based on your experience running ApowerB:

On observability and trace storage: Do the bundled th2pulse and otel-collector integrate nicely if an engineering team already runs their own centralized stack (like Grafana Tempo / Prometheus Mimir), or is it better to just leave them running locally within the chart?

On agent pools: For heavier agent workloads or resource-intensive runs, would you recommend switching early to external S3 object storage instead of the PVC, or does local PVC storage hold up fine for initial production traffic?

Great work breaking down the real-world operational side, especially calling out the WaitForFirstConsumer trap! 🚀

Collapse
 
anis_meziani_52aab42304a8 profile image
Anis Meziani •

Thanksfor ur comment. On observability: the otel-collector is made for exactly that, point its exporters at your Tempo/Prometheus/Mimir and forward from day one. th2pulse is just the local default for teams without a stack.

On storage: switch to S3 earlier than you'd think, and the trigger is access mode, not size. The PVC is ReadWriteOnce, so the moment you scale the backend past one replica or share artifacts across pods, that's the bottleneck, not capacity. Single-replica pilot, PVC is fine.