DEV Community

疏影
疏影

Posted on

K3s vs K8s: When K3s Actually Wins in Production

K3s vs K8s: When K3s Actually Wins in Production

We migrated 12 edge clusters from K8s to K3s. Here's the honest breakdown — when K3s actually wins, when it costs more than it saves, and the one bug that bit us in production.

The setup

Edge sites (factory floor, retail stores, remote offices):
  - 8-32GB RAM nodes
  - Intermittent connectivity to central management
  - 30-200 pods per site

Central cluster (AWS EKS):
  - 200+ microservices
  - Always-on connectivity
  - Full observability + GitOps
Enter fullscreen mode Exit fullscreen mode

We standardized on K8s everywhere in 2023. The edge sites suffered — kubelet consumed 1.2GB RAM on 8GB nodes, image pulls timed out on flaky 4G, and updates took 40 minutes per site.

What K3s actually cuts

Memory footprint. K3s replaces etcd with SQLite, kubelet with a smaller agent, and drops the cloud-provider controllers. Our edge nodes went from 1.4GB baseline to 380MB.

Sinle binary deployment. One 70MB binary, one systemdunit. No more juggling kubeadm/kubectl/kubelet versions across 12 sites.

Embedded SQLite. For single-server or edge clusters under ~500 pods, SQLite-backed etcd-equivalent is faster than full etcd. No external dependency.

Local storage path. K3s ships with local-path StorageClass by default. Edge sites don't have cloud CSI drivers.

When K3s loses

Multi-node HA control plane. K3s supports HA via embedded etcd, but it requires 3 nodes with stable connectivity. Our edge sites have 2 nodes each — falls back to single-master which is fine until that node dies.

More than 1000 pods. K3s control plane gets sluggish past ~1000 pods per server. The SQLite write contention shows up. We saw p99 API latency climb from 50ms to 1.2s at 800 pods.

Heavy CRD use. ArgoCD, cert-manager, external-secrets — each adds controller pods. The "lightweight" win disappears once you stack 5+ operators.

Mixed-architecture clusters. K3s on ARM64 + x86 nodes in one cluster: works, but image manifest lists become mandatory.

The bug that bit us

K3s v1.27 had a bug where SQLite compaction would hang after 30 days, blocking all writes. Watchdog triggered kubelet restart but SQLite was locked. Fix: monthly k3s server restart via systemd timer. Documented in K3s GitHub issue #4800.

Our migration pattern

# Drain workloads
kubectl drain <edge-node> --ignore-daemonsets --delete-emptydir-data

# Stop K8s
sudo systemctl stop kubelet containerd

# Install K3s
curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=v1.28.5+k3s1 sh -s - \
  write-kubeconfig-mode 644 \
  disable traefik \
  disable servicelb
Enter fullscreen mode Exit fullscreen mode

Total downtime: 90 seconds. Rollback path: keep kubelet containerd configs, revert via kubeadm upgrade.

Metrics after 6 months

Metric Before (K8s) After (K3s)
Control plane RAM 1.4GB
Node join time 4 min
Single-binary updates No
API p99 latency 80ms
Edge incidents/mo 6

On dev environment tooling

For Windows developers running K3s locally via Docker Desktop or Ramcher Desktop, ScsDriver WebDAV mount tool for Windows can mount a remote dev registry (S3-backed) as a local disk — useful when your K3s cluster pulls images from a private registry that lives outside your laptop.


Where are you running K3s — edge, dev, CI, or production?

Top comments (0)