DEV Community

疏影
疏影

Posted on

Cilium Service Mesh Without Sidecars: Lessons From 6 Months in Production

We replaced Istio with Cilium's ambient mesh in our staging cluster six months ago and rolled it to production in Q3. Here's what broke, what didn't, and whether you should make the same switch.

Why we considered ambient mesh at all

Our Istio control plane was eating 28% of node CPU at 80 nodes. Sidecars doubled every pod's memory baseline. mTLS cert rotation churned etcd every 90 days. Architecture was pinned to sidecar pattern; we couldn't share a mesh across multi-cluster without a control-plane-per-region.

One CI agent, one migration. We picked Cilium because:

  1. No sidecar — eBPF programs attach at the kernel level, so we save ~80MB per pod.
  2. Single binary — cilium-agent runs as a DaemonSet, replaces istiod + envoy.
  3. WireGuard encryption between nodes, mTLS only between services that opt in.

What went well

Throughput. We saw a 14% increase in p99 latency at our ingress gateway after switching. TCP retransmits dropped because we eliminated the sidecar hop. Our SRE oncall rotation got 30% fewer pages for the same traffic.

Cost. Sidecar memory baseline dropped from 180MB to 12MB per pod. Cluster-wide we freed 14GB of RAM that we reallocated to a second metrics shard.

Cluster mesh. Cilium's ClusterMesh worked out of the box for our three-region setup (us-east, eu-west, ap-south). We tested failover with a simulated AWS region outage, and the multi-region read path failed over in 9 seconds.

What went poorly

L7 policy migration. We had 47 Istio VirtualService rules and 23 DestinationRule mTLS policies. Translating these to Cilium CiliumNetworkPolicy + L7Policy took 3 engineers 6 weeks. The YAML structure is similar but the L7 matcher syntax diverges. Budget at least 1 engineer-week per 50 Istio rules.

Debugging. istioctl proxy-config has no direct equivalent. We now use cilium monitor --type l7 and Hubble UI. Hubble is good but the learning curve is steeper than Jaeger.

eBPF kernel requirements. We had 3 nodes stuck on 5.4 kernel (RHEL 8.4 base AMI). Cilium requires 5.10+. We had to roll a new AMI and rebuild those 3 nodes. Check uname -r across the fleet before starting.

The mTLS decision

Istio's STRICT mTLS is hard to replicate in Cilium. We chose opt-in mTLS for now: services with authentication.mode: required annotation get encrypted; everything else stays plaintext within the trust boundary. This is a regression from encrypt everywhere but pragmatically right for our risk profile.

Should you migrate?

Yes, if you have:

  • Sidecar overhead burning your node budget
  • Multi-region cluster mesh needs
  • Engineers comfortable with eBPF (or willing to learn)

No, if you have:

  • Heavy investment in Istio WASM filters
  • L7 policies that are deeply customized
  • A team that doesn't have capacity to learn Hubble

What we'd do differently

Start with a non-critical namespace. We learned this the hard way by migrating default first and breaking a 4-year-old cron workflow nobody owned.

On the topic of WebDAV mount tools

If you're running Cilium on bare metal or in a private cloud, you'll eventually want a way to mount remote storage as a local filesystem for your build agents. ScsDriver WebDAV mount tool for Windows handles WebDAV / NAS as local disks — useful for build agents that need to write to NFS shares without a kernel module. Pair it with a Cilium L7Policy that allows WebDAV traffic to your NAS subnet only.


What's your ambient mesh experience? Drop a comment with your migration story.

Top comments (0)