Hermes Agent is Nous Research's self-improving AI agent. It builds skills from experience, keeps persistent memory across sessions, and connects to Telegram, Discord, Slack, WhatsApp, and Signal out of the box. With 230k+ GitHub stars, plenty of teams now want to run it somewhere more durable than a laptop.
Kubernetes is the obvious destination and the least documented one. There's no official Helm chart. The container image's init sequence needs root in a way that collides with a standard restricted pod security context. Its state model assumes a single writer. And the reload story that works fine in a terminal has no equivalent your GitOps controller can call.
None of that is a reason to avoid Hermes on Kubernetes. It's a reason to know the constraints before you write the manifest. Below: what genuinely breaks, what only looks like it breaks, and a reference deployment to start from.
Tested against: Hermes Agent
v2026.8.27· official imagenousresearch/hermes-agent· Kubernetes 1.31+ · containerd 2.x. Hermes ships releases every few days. Before you copy anything below, check it against the version you're actually deploying, and pin that version.
Is there an official Helm chart for Hermes Agent?
No. Nous Research's own documentation covers install.sh, Docker and Docker Compose, and Nix packages. Kubernetes never comes up as a deployment target. That gap is filled entirely by the community, and the three charts that fill it don't agree on much:
-
ultraworkers/hermes-agent-helm-chart: the most feature-complete option: renders a Deployment, PVC, Secret, Service, Ingress, and an Istio VirtualService, plus an "operator-ready" mode defining a
HermesTenantCRD. It's also the only one that encodes the single-writer rule as a hard constraint. - jyje/hermes-agent: listed on Artifact Hub as a verified publisher, distributed over an OCI registry, with multi-arch images and example ArgoCD manifests. It tracks upstream Hermes releases more closely than the other two.
-
duyet/hermes-agent: the simplest of the three, splitting persistence into separate data and workspace volumes, with an optional Prometheus
ServiceMonitor. Not a verified publisher, so read the templates before trusting the defaults.
None of the three is backed by Nous Research. Read the templates end to end before you apply one, and expect to override the security context.
Hermes Agent is stateful: define your persistence boundary first
This is the decision everything else hangs off, so make it before you pick a chart.
Hermes keeps all of its mutable state (configuration, MEMORY.md, USER.md, session history in a local SQLite database, and every learned skill) under HERMES_HOME. In the official image that's the /opt/data volume, which the Docker docs call "the single source of truth for all Hermes state."
So the accurate invariant is not "Hermes can't scale." It's:
Treat each
HERMES_HOMEas a single-writer state domain.
You can run many Hermes instances in one cluster. What you must not do is put two active pods behind the same mutable state and assume Kubernetes has handed you horizontal scaling. Nothing arbitrates concurrent writes to that SQLite database and those Markdown files.
In practice, for any deployment with persistence enabled:
-
replicaCountstays at 1. -
strategy.typemust beRecreate, notRollingUpdate, a rolling update deliberately runs the old and new pod together, which is exactly the two-writer window to avoid. -
ReadWriteOnceis the sensible guardrail on the PVC. RWX isn't automatically unsafe (the invariant is one active writer, not one mount) but RWO enforces it at the storage layer instead of trusting your deployment config.
Need Hermes for more than one team or tenant? One release per tenant, each with its own volume. Any chart that doesn't enforce this is a data-corruption risk waiting for a bad rollout.
Why the official image fights runAsNonRoot
The common shorthand, "Hermes has to run as root," is wrong, and the precise version matters when you're arguing with a platform team about a Pod Security Standard exemption.
The official image uses s6-overlay as its init system. s6-overlay's /init runs as root so it can chown the volume on first boot, then drops to the hermes user via s6-setuidgid for the main program and all supervised services.
So the agent process itself does not run as root. Only the bootstrap does. But that's still enough to break a restricted security context: a pod spec with runAsNonRoot: true or a non-zero runAsUser blocks the sequence before the agent ever starts, and you get something close to:
/package/admin/s6-overlay/libexec/preinit: fatal: /run belongs to uid 0 instead of 1000,
has insecure and/or unworkable permissions, and we're lacking the privileges to fix it.
s6-overlay-suexec: fatal: child failed with exit code 100
One detail worth getting right, because published examples routinely get it wrong: the hermes user is UID 10000, not 1000. Set fsGroup: 1000 and the volume ends up owned by the wrong group.
A reasonable starting security context:
podSecurityContext:
runAsUser: 0
runAsNonRoot: false
fsGroup: 10000
fsGroupChangePolicy: OnRootMismatch
seccompProfile:
type: RuntimeDefault
securityContext:
allowPrivilegeEscalation: true
readOnlyRootFilesystem: false
runAsNonRoot: false
capabilities:
drop: ["ALL"]
add: ["CHOWN", "SETUID", "SETGID"]
That is meaningfully narrower than a privileged pod: every capability is dropped and only the ones the bootstrap needs come back. Treat the exact capability list as a starting point to verify against your image version, not a universal recipe. Start from drop: ["ALL"], add back only what your logs prove necessary, and re-test on version bumps.
If your cluster enforces the restricted Pod Security Standard, this workload needs a namespace exemption. Forcing the official image to start non-root isn't a setting you've missed, it wasn't built for that.
What Hermes can hot-reload, and what it can't
The claim that Hermes has "no hot-reload" is false, and the real limitation is more interesting.
Hermes documents three reload commands: /reload-mcp (reload MCP servers from config.yaml), /reload-skills (re-scan for newly installed or removed skills), and /reload (reload .env variables into the running session).
The runtime can absolutely pick up new MCP servers and skills without a restart. The gap is the interface:
Hermes can hot-reload MCP servers and skills interactively, but doesn't expose the same lifecycle through its admin API.
That's the actual Kubernetes problem:
ConfigMap updated -> kubelet syncs file into the pod -> no reconciliation hook -> nothing happens
The file changes on disk. Nothing tells the process. Your options:
-
Restart the workload.
kubectl rollout restart deployment/hermes-agent. Blunt but declarative, and withRecreateyou're accepting a brief outage anyway. Wire a checksum annotation over the ConfigMap into the pod template so the rollout fires automatically on config change, the standard Helmchecksum/configpattern. - Drive the runtime reload path from a sidecar or an operator with a channel into the agent. Workable, but you're building the missing hook yourself.
Until an HTTP reload endpoint exists upstream, option 1 with a config checksum is the pragmatic default.
Resources, probes, and graceful shutdown
Three things almost every Hermes-on-Kubernetes writeup omits.
Resources. Nous Research's Docker documentation recommends a minimum of 1 GB memory and 1 CPU core, with 2-4 GB memory and 2 cores recommended. Browser automation (Playwright/Chromium) is the most memory-hungry feature, so budget above that range if the agent drives a browser. Set a memory limit, but consider leaving CPU unlimited, agent workloads are bursty, and CPU throttling shows up as mysteriously slow tool calls rather than a clean failure.
Probes. The gateway listens on port 8642. The Dockerfile ships no HEALTHCHECK and no EXPOSE instruction, so you're defining this yourself. A tcpSocket probe against 8642 is the portable choice. The one that matters most is the startupProbe: first boot does volume chown work and profile reconciliation, and a liveness probe with a short threshold will kill the pod mid-bootstrap and loop forever.
Graceful shutdown. Hermes is stateful, so SIGTERM handling isn't academic. Give it room with terminationGracePeriodSeconds: 30, and test what actually happens to in-flight agent runs, SQLite session writes, and open gateway connections when a pod is evicted. Node upgrades and spot reclaims will do this to you eventually, and Recreate means no second pod is covering the gap.
Two security boundaries, not one
Most Kubernetes writeups on agents cover pod security and stop. For an agent that executes code, that's half the problem.
Boundary 1, pod security. Root init, capabilities, seccomp, filesystem, service account. Covered above.
Boundary 2, agent execution. What the model can actually run, and what it can reach. This is the one that should worry you more, because the pod is the blast radius.
Hermes has real defences here:
-
Credential filtering.
execute_codeblocks environment variables whose names containKEY,TOKEN,SECRET,PASSWORD,CREDENTIAL,PASSWD, orAUTH. MCP stdio subprocesses receive only a short safe list of variables. -
The bypass is explicit. Variables declared by a skill or listed in
env_passthroughskip those filters. That's the mechanism to audit. -
Approval checks are skipped in container backends. Hermes skips dangerous-command approval in the
docker,singularity,modal,daytona, andvercel_sandboxbackends, on the reasoning that "the container itself is the security boundary." In a Kubernetes pod that assumption is load-bearing: whatever your pod can reach, prompt-injected code can reach. - SSRF protection exists, and is disableable. Web tools block RFC 1918 ranges and loopback by default. In a cluster, RFC 1918 is your service mesh, your databases, and the kubelet.
Which leads to the control most often missing: default-deny egress. An agent that browses the web and calls tools is not a normal web app, and "what can this pod reach?" is a more consequential question than which Linux capabilities it holds. Start from deny-all and allow only what's needed:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: hermes-agent-egress
spec:
podSelector:
matchLabels:
app.kubernetes.io/name: hermes-agent
policyTypes: ["Egress"]
egress:
# DNS
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: kube-system
podSelector:
matchLabels:
k8s-app: kube-dns
ports:
- protocol: UDP
port: 53
# HTTPS to the internet, minus internal ranges
- to:
- ipBlock:
cidr: 0.0.0.0/0
except:
- 10.0.0.0/8
- 172.16.0.0/12
- 192.168.0.0/16
- 169.254.169.254/32 # cloud instance metadata
ports:
- protocol: TCP
port: 443
Blocking 169.254.169.254 matters especially: on a node without IMDSv2 enforced, an agent that can reach instance metadata can often reach the node's IAM role.
On credentials: because the agent has terminal access, anything in its environment is potentially readable by it. On EKS, use IRSA or EKS Pod Identity rather than static access keys in a Secret. Elsewhere, External Secrets Operator, Vault, or Sealed Secrets all keep plaintext keys out of Git. And scope the IAM role tightly: the agent's permissions are the agent's capabilities.
A minimal deployment that actually works
Pinned, single-writer, probed, and resource-bounded. Adjust the storage class and image tag, and read it before you apply it.
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: hermes-agent-data
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 20Gi
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: hermes-agent
labels:
app.kubernetes.io/name: hermes-agent
spec:
replicas: 1
strategy:
type: Recreate # never two writers on one HERMES_HOME
selector:
matchLabels:
app.kubernetes.io/name: hermes-agent
template:
metadata:
labels:
app.kubernetes.io/name: hermes-agent
annotations:
checksum/config: "REPLACE_WITH_CONFIGMAP_CHECKSUM"
spec:
serviceAccountName: hermes-agent
terminationGracePeriodSeconds: 30
securityContext:
runAsUser: 0
runAsNonRoot: false
fsGroup: 10000
fsGroupChangePolicy: OnRootMismatch
seccompProfile:
type: RuntimeDefault
containers:
- name: hermes-agent
image: nousresearch/hermes-agent:v2026.8.27 # pin it; never :latest
args: ["gateway", "run"]
ports:
- name: gateway
containerPort: 8642
envFrom:
- secretRef:
name: hermes-agent-secrets
securityContext:
allowPrivilegeEscalation: true
readOnlyRootFilesystem: false
runAsNonRoot: false
capabilities:
drop: ["ALL"]
add: ["CHOWN", "SETUID", "SETGID"]
resources:
requests:
cpu: "500m"
memory: 2Gi
limits:
memory: 4Gi
startupProbe:
tcpSocket:
port: gateway
periodSeconds: 5
failureThreshold: 60
readinessProbe:
tcpSocket:
port: gateway
periodSeconds: 10
livenessProbe:
tcpSocket:
port: gateway
periodSeconds: 20
failureThreshold: 3
volumeMounts:
- name: data
mountPath: /opt/data
volumes:
- name: data
persistentVolumeClaim:
claimName: hermes-agent-data
---
apiVersion: v1
kind: Service
metadata:
name: hermes-agent
spec:
selector:
app.kubernetes.io/name: hermes-agent
ports:
- name: gateway
port: 8642
targetPort: gateway
Pair it with the NetworkPolicy above, a Secret (ideally rendered by External Secrets or Sealed Secrets rather than committed), and a ConfigMap for config.yaml if you're managing MCP servers declaratively.
Production checklist
- Image tag pinned to a specific release, never
:latest -
replicas: 1andstrategy.type: Recreate, with oneHERMES_HOMEper instance - PVC is
ReadWriteOnce, and the storage class supports that access mode -
fsGroup: 10000matches the image'shermesuser - Capabilities start from
drop: ["ALL"], with additions verified against your image version - Namespace exemption in place if you enforce the
restrictedPod Security Standard -
startupProbegenerous enough to survive first-boot volume work -
terminationGracePeriodSecondsset, and eviction behaviour tested against in-flight runs - Memory limit sized for browser automation if skills use it
- Default-deny egress NetworkPolicy, with instance metadata (
169.254.169.254) blocked - No static cloud credentials in Secrets, use IRSA / Pod Identity / Vault / External Secrets
-
env_passthroughand skill-declared variables audited, they bypass credential filtering - Config changes trigger a rollout (checksum annotation) or a deliberate reload path
Where this connects to governed data
Everything above keeps Hermes Agent alive and contained. None of it says whether the answers it produces are correct, that depends on what it's allowed to read and how well-defined that data is. The same discipline that makes an agent's infrastructure trustworthy (scoped permissions, no unmanaged state, changes that go through review) reappears in agentic data engineering, where an agent's output earns trust through layered checks rather than raw model capability. Point an agent like this at real business data instead of chat platforms and a semantic layer is what stops it guessing at what "revenue" or "active customer" means.
This article was originally published on the RevOS blog.
Top comments (1)
The "define your persistence boundary first" advice is the part people skip and then pay for. Single-writer state on Kubernetes is deceptively dangerous because the platform's whole instinct — rolling updates, multiple replicas, reschedule-on-drain — is built around the assumption that your workload isn't a single writer. A plain
Deploymentwill happily run the old and new pod simultaneously during a rollout, and that overlap window is exactly when a single-writer state model corrupts.The usual fix is a
StatefulSetwith.spec.updateStrategyand aRollingUpdatepartition, or justRecreatesemantics if you can tolerate the brief downtime — anything that guarantees the old writer is fully gone before the new one touches the volume. TheReadWriteOncePVC helps but doesn't save you across a node failover if the old pod is only network-partitioned, not dead.Did the community charts handle the terminationGracePeriod / preStop drain cleanly, or did you have to add your own hook to make sure the writer flushes before the volume detaches? That handoff is where I'd expect the subtle data loss to hide.