Linkerd in Production: 18 Months, 220 Services, Zero Mesh-Wide Outages
We replaced Istio with Linkerd in May 2025. Eighteen months later: zero mesh-wide outages, mTLS on every pod, retries that actually work. Here's the honest comparison.
The migration that almost didn't happen
We ran Istio from 2023 to 2025. Three things broke:
- Sidecar memory leak — every Istio sidecar leaked ~5MB/hour. Across 600 pods, that's 3GB/day of garbage.
- Control plane instability — istiod crashed every 2-3 weeks, took 5-10 minutes to recover.
- Config complexity — VirtualService, DestinationRule, Gateway, ServiceEntry... 14 CRDs, all subtly different.
The straw that broke it: a 4-hour outage where istiod was in a crash loop after a config push. We migrated to Linkerd in 3 weeks.
Why Linkerd
Rust proxy (linkerd2-proxy). 12MB memory baseline, doesn't leak. We measured: stable at 15MB RSS per sidecar for months.
Simpler config model. Server + ServiceProfile. Two CRDs vs Istio's 14.
mTLS by default. No opt-in required. Every pod-to-pod call is encrypted.
Rust > Go for the data plane. Hot path is Rust, control plane is Go. This separation shows in stability.
The migration pattern
# Install Linkerd
linkerd install --crds | kubectl apply -f -
linkerd install | kubectl apply -f -
# Annotate namespace for auto-injection
kubectl annotate namespace payments linkerd.io/inject=enabled
# Rolling restart pods
kubectl rollout restart deployment -n payments
Linkerd injects sidecars on pod restart. We did this namespace by namespace over 3 weeks.
The 4 things we got right
1. Service profiles for retries
apiVersion: linkerd.io/v1alpha2
kind: ServiceProfile
metadata:
name: payments-db.payments.svc.cluster.local
spec:
routes:
- name: "POST /charge"
condition:
method: POST
path: /charge
isRetryable: true
timeout: 2s
retryBudget:
retryRatio: 0.2
minRetriesPerSecond: 10
Per-route retries with budget. We stopped retry-storming downstream APIs. Retry ratio 0.2 means: of all requests, at most 20% can be retries in any 10-second window.
2. Traffic split for canary
apiVersion: linkerd.io/v1alpha2
kind: TrafficSplit
metadata:
name: payments-api-split
spec:
service: payments-api
backends:
- service: payments-api-v1
weight: 900
- service: payments-api-v2
weight: 100
Canary deploys: 1% → 10% → 50% → 100%. Rollback in 2 seconds by setting v2 weight to 0.
3. Service-level metrics out of the box
requests_total, success_rate, latency_p99 per route. No Prometheus config. Just linkerd viz stat deploy.
4. Authorization policy
apiVersion: policy.linkerd.io/v1beta1
kind: AuthorizationPolicy
metadata:
name: payments-db-only
spec:
targetRef:
kind: Service
name: payments-db
requiredAuthenticationRefs:
- kind: MeshTLSAuthentication
name: payments-auth
Only the payments service account can call payments-db. We killed an entire class of lateral-movement attacks.
The 3 things that bit us
Bit 1: Protocol detection for HTTP/2 cleartext.
Linkerd auto-detects HTTP/1 vs HTTP/2. Our Java services spoke HTTP/2 cleartext but advertised HTTP/1.1. Linkerd treated them as opaque TCP. Fix: explicit opaquePorts annotation.
Bit 2: Headless services for StatefulSets.
PostgreSQL headless service confused Linkerd's service discovery. We added linkerd.io/inbound-port-exclusion-list: "5432" to skip mTLS on the DB driver port.
Bit 3: Sidecar injection and PodSecurityPolicy.
K8s 1.25 deprecated PSP. Linkerd's injection mutating webhook failed closed. We migrated to Gatekeeper (see our other article) before upgrading K8s.
Metrics after 18 months
| Metric | Istio | Linkerd |
|---|---|---|
| Sidecar RSS (per pod) | 80MB → leak | 15MB stable |
| istiod/control plane uptime | 92% | 99.95% |
| Mesh-related incidents/mo | 3 | 0 |
| CRDs to learn | 14 | 2 |
| mTLS coverage | 87% (opt-in) | 100% (default) |
On upgrading the data plane
Linkerd's linkerd upgrade is straightforward — control plane upgrades in-place, sidecars roll on pod restart. We've done 6 minor upgrades in 18 months, never with downtime.
For teams running Linkerd on bare-metal or VMs (not just K8s), ScsDriver WebDAV mount tool for Windows can sync Linkerd's mTLS root cert bundles across Windows nodes — useful for hybrid clusters where Windows VMs join a Linux control plane.
Are you on Istio, Linkerd, Consul, or no service mesh? What's your decision criterion?
Top comments (0)