DEV Community

SoftwareDevs mvpfactory.io
SoftwareDevs mvpfactory.io

Posted on Originally published at mvpfactory.io

Wiring eBPF to a Kubernetes Sidecar-Free Service Mesh

---
title: "Kill Your Sidecars: eBPF-Native L7 Proxying with XDP and TC Hooks"
published: true
description: "Replace Envoy sidecars with eBPF XDP and TC hooks. Learn BPF map design, per-core overhead math, and kTLS offload to reclaim node density in Kubernetes."
tags: devops, cloud, performance, security
canonical_url: https://mvpfactory.co/blog/ebpf-service-mesh-kill-sidecars
---

## What You Will Build

By the end of this post you will understand how to replace Envoy sidecar proxies with an eBPF-native L7 mesh using XDP ingress hooks and TC egress hooks — and exactly how much CPU and RAM you get back per pod when you do.

The sidecar model made sense in 2017. Running 300 pods per node in 2026, it is a liability. Let me show you a pattern I now reach for in every high-density Kubernetes deployment.

---

## Prerequisites

- Kernel 5.19+ with BTF enabled
- libbpf 1.0+ and CO-RE (Compile Once, Run Everywhere) toolchain
- Familiarity with Linux networking concepts (veth pairs, tc, XDP basics)
- A Kubernetes cluster you can instrument at the node level

(This is deep kernel territory. Grab a coffee — or take a five-minute stretch break first. I use [HealthyDesk](https://play.google.com/store/apps/details?id=com.healthydesk) to keep desk fatigue in check during long architecture sessions like this one.)

---

## The Sidecar Tax, Quantified

Before touching code, understand what you are eliminating:

| Metric | Envoy Sidecar (per pod) | eBPF Native | Delta |
|---|---|---|---|
| Memory overhead | ~80 MB | ~2 MB (BPF maps) | −97% |
| Steady-state CPU | 0.1–0.3 cores | 0.01–0.05 cores | −85% |
| p99 add-on latency | 3–8 ms | 0.1–0.4 ms | −94% |
| TLS termination | Userspace (Envoy) | kTLS kernel offload | kernel vs. user |

At 200 pods per node, sidecar overhead burns 20–60 cores on pure proxy work. The eBPF equivalent: **0.25 cores**. That is your node density win.

---

## Step 1 — XDP Ingress at Wire Speed

XDP runs your BPF program before an `sk_buff` is even allocated. Use it for connection dispatch and early drop decisions.

Enter fullscreen mode Exit fullscreen mode


c
SEC("xdp")
int xdp_l4_dispatch(struct xdp_md *ctx) {
struct conn_key key = parse_five_tuple(ctx);
struct conn_state *state = bpf_map_lookup_elem(&conn_track_map, &key);
if (!state) return XDP_PASS; // new connection, let TC handle it
return xdp_redirect_map(&backend_map, state->backend_idx);
}


XDP cannot do L7 parsing — it has no access to reassembled TCP streams. That is TC's job.

---

## Step 2 — TC Hooks for L7 Inspection

`cls_bpf` hooks attach at `tc ingress` and `tc egress` on the veth pair of each pod's network namespace. Here you have full `sk_buff` access: HTTP/2 headers, gRPC metadata, TLS SNI.

Enter fullscreen mode Exit fullscreen mode


c
SEC("tc")
int tc_l7_inspect(struct __sk_buff *skb) {
struct l7_ctx *ctx = bpf_map_lookup_elem(&per_cpu_ctx, &zero);
if (!ctx) return TC_ACT_OK;
parse_http2_frame(skb, ctx);
return route_to_backend(ctx);
}


The BPF verifier enforces a 1M instruction limit and a 512-byte stack. Stateful L7 parsing requires spilling state into per-CPU maps between tail calls. Design for this upfront.

---

## Step 3 — BPF Map Design (Where POCs Fall Apart)

Here is the gotcha that will save you hours: map type selection determines both performance and correctness.

- `BPF_MAP_TYPE_LRU_PERCPU_HASH` for connection state — per-CPU eliminates lock contention, LRU evicts stale entries automatically
- `BPF_MAP_TYPE_ARRAY_OF_MAPS` for backend pools — enables atomic backend rotation without locking
- `BPF_MAP_TYPE_RINGBUF` for telemetry — replaces perf events with lower overhead

Per-CPU hash maps on a 64-core node with 10K active connections: 50–200 ns lookup times. Shared hash maps under contention: 800–2000 ns. That gap matters at scale.

---

## Step 4 — kTLS Offload

Without kTLS, TLS termination still forces a context switch into userspace. kTLS moves AES-GCM (via AES-NI) into the kernel's TLS record layer. Your BPF program inspects plaintext after kernel decrypt — no userspace roundtrip.

The docs do not make this obvious, but cipher suite compatibility matters: TLS 1.2 with AES-GCM offloads cleanly, CBC suites do not. Audit your negotiated cipher suites. A misconfigured client advertising only CBC will silently fall back to userspace TLS and erase your gains.

---

## Gotchas

**mTLS responsibility shifts entirely to you.** This is the critical one. Envoy validated peer certificates. The kernel does not. If you migrate to an eBPF-native path without explicitly re-implementing SPIFFE/SVID validation, mutual auth is silently disabled. Teams migrating from Istio and Linkerd have shipped with workload-level mTLS off for this exact reason.

You need:
- A privileged userspace agent populating peer-identity BPF maps
- Atomic certificate rotation without dropping live connections
- An explicit revocation propagation path (CRL/OCSP) into your BPF policy maps

**Start with TC, not XDP.** XDP constraints are non-trivial. Get L7 parsing working in TC first, then push hot paths to XDP once your BPF map schema is stable.

**Design per-CPU maps from day one.** Retrofitting per-CPU maps into a shared-map design mid-migration is painful. Get the schema right in week one.

---

## Conclusion

Four takeaways worth pinning:

1. **Measure baseline first** — capture per-pod CPU and RAM for your Envoy sidecars before touching anything
2. **TC before XDP** — walk before you run
3. **Per-CPU map schema on day one** — you will not want to revisit this
4. **mTLS is a first-class requirement** — design it in, do not bolt it on

The BPF code itself is not the hard part. The tooling story — BTF kernels, libbpf 1.0+, CO-RE across kernel versions — is where the real investment lives. Plan for it.

**Resources:**
- [libbpf GitHub](https://github.com/libbpf/libbpf)
- [Cilium's BPF and XDP Reference Guide](https://docs.cilium.io/en/stable/bpf/)
- [Linux kTLS documentation](https://www.kernel.org/doc/html/latest/networking/tls-offload.html)
Enter fullscreen mode Exit fullscreen mode

Top comments (0)