How eBPF Actually Runs Inside the Linux Kernel: The Verifier, JIT, and BPF Maps
If you ask most developers how eBPF works, you usually get one of two answers:
- "It is a tiny virtual machine inside Linux, like a JVM for the kernel."
- "It is a way to run custom C code in the kernel without writing a kernel module."
Both explanations miss the fundamental architecture.
eBPF does not run inside an interpreter loop at runtime, and it definitely does not allow arbitrary C execution. If you could execute arbitrary C in kernel space, a single null pointer dereference or off-by-one loop would panic your entire operating system.
Instead, eBPF is an in-kernel sandboxed execution engine governed by an unforgiving static verifier, compiled to native machine instructions via a Just-In-Time (JIT) compiler, and bridged to user space through lockless, memory-mapped data structures.
Here is the exact mechanical journey of an eBPF program: from user-space bytecode compilation, through the kernel's symbolic execution verifier, down to native CPU execution on kernel hooks.
1. The Virtual CPU: Registers and Calling Convention
Before touching the kernel, your high-level C or Rust code is compiled by LLVM/Clang targeting the bpf backend (clang -target bpf -O2).
The output is not x86 or ARM assembly. It is 64-bit eBPF bytecode designed around an idealized RISC register machine:
+-----------------------------------------------------------------------+
| eBPF Register Architecture |
+-----------------------------------------------------------------------+
| Register | Purpose |
+------------+----------------------------------------------------------+
| R0 | Function return value & exit code to kernel hook |
| R1 - R5 | Function arguments passed to in-kernel BPF helpers |
| R6 - R9 | Callee-saved registers (preserved across helper calls) |
| R10 | Read-only frame pointer to 512-byte stack frame |
+-----------------------------------------------------------------------+
Every eBPF instruction is fixed at 64 bits (8 bytes):
- 8-bit opcode: Operation type (ALU, memory load/store, branch jump).
-
4-bit destination register (
dst_reg): Target register (0–10). -
4-bit source register (
src_reg): Source register (0–10). - 16-bit offset: Signed offset for stack or memory addressing.
-
32-bit immediate (
imm): Constant integer operand.
0 7 8 11 12 15 16 31 32 63
+---------+-------+-------+-------------------+-------------------------+
| Opcode | Dst | Src | Offset | Immediate |
+---------+-------+-------+-------------------+-------------------------+
Unlike user-space programs that have megabytes of stack space, each eBPF program gets exactly 512 bytes of stack pointed to by R10. Attempting to allocate a large struct on the stack will fail before your code ever runs.
2. Loading Bytecode: The bpf() Syscall
When a user-space loader (written in Go with cilium/ebpf, Rust with aya, or C with libbpf) wants to load an eBPF program, it invokes the bpf() system call with the BPF_PROG_LOAD command:
union bpf_attr attr = {
.prog_type = BPF_PROG_TYPE_KPROBE,
.insns = (uint64_t)bytecode_buffer,
.insn_cnt = instruction_count,
.license = (uint64_t)"GPL",
.log_buf = (uint64_t)verifier_log_buffer,
.log_size = 65536,
.log_level = 1,
};
int prog_fd = syscall(SYS_bpf, BPF_PROG_LOAD, &attr, sizeof(attr));
The kernel receives this bytecode buffer, but it does not execute it immediately. Instead, it hands the raw instruction stream over to kernel/bpf/verifier.c.
3. The Verifier: Symbolic Execution and Safety Proofs
The verifier is the security heart of eBPF. Its job is mathematically proving that your program cannot crash the kernel, dereference null pointers, read uninitialized memory, or loop infinitely.
The verification process runs in two distinct phases:
[ Bytecode Input ]
│
▼
┌────────────────────────────────────────┐
│ Phase 1: DAG & CFG Validation │
│ - Detect unreachable instructions │
│ - Check for uncontrolled loops / cycles│
└──────────────────┬─────────────────────┘
│ Pass
▼
┌────────────────────────────────────────┐
│ Phase 2: Abstract Interpretation │
│ - Simulate every execution path │
│ - Track register states (R0-R10) │
│ - Validate pointer types and offsets │
│ - Enforce bounds & NULL checks │
└──────────────────┬─────────────────────┘
│ Pass
▼
[ JIT Compiler -> Native Machine Code ]
Phase 1: Control Flow Graph (CFG) Validation
The verifier parses the program into a Directed Acyclic Graph (DAG). It ensures:
- The program has a clean entry point at instruction 0 and ends on a valid
BPF_EXIT_INSN. - There is no unreachable dead code.
- Loops are strictly bounded. Historically, all loops were forbidden. Modern kernels allow bounded loops, but the verifier must prove that the loop terminates within a fixed instruction budget (up to 1,000,000 processed instructions during simulation).
Phase 2: Abstract Interpretation (Register State Tracking)
Next, the verifier descends through every possible execution branch, simulating CPU execution instruction by instruction.
For every step, the verifier maintains a struct bpf_reg_state for all 11 registers. A register is never just a raw number; the kernel assigns it a strict type:
-
NOT_INIT: Register has not been written to yet. Reading it is an immediate failure. -
SCALAR_VALUE: An integer whose possible minimum and maximum values are tracked viatnum(tristate numbers tracking known 0s, known 1s, and unknown bits). -
PTR_TO_CTX: Pointer to the hook's context struct (for example,struct xdp_mdorstruct pt_regs). -
PTR_TO_MAP_VALUE_OR_NULL: Pointer returned by a map lookup that might be valid or might be0(NULL). -
PTR_TO_MAP_VALUE: A verified, non-null pointer to map memory. -
PTR_TO_STACK: A pointer to the local 512-byte stack frame.
Why the Verifier Rejects Unchecked Pointers
Look at this common bug in C:
struct packet_stat *stat = bpf_map_lookup_elem(&stats_map, &key);
stat->packets += 1; // Crash if key does not exist!
When compiled to eBPF bytecode and evaluated:
-
bpf_map_lookup_elem()executes. The verifier marksR0with the typePTR_TO_MAP_VALUE_OR_NULL. - The next instruction attempts a memory store:
*(u64 *)(R0 + 0) += 1. - The verifier checks the type of
R0. BecauseR0is stillPTR_TO_MAP_VALUE_OR_NULL(and notPTR_TO_MAP_VALUE), the verifier aborts with:
R0 invalid mem access 'map_value_or_null'
To pass verification, you must write a conditional branch:
struct packet_stat *stat = bpf_map_lookup_elem(&stats_map, &key);
if (!stat) {
return 0; // Branch A: R0 is NULL, exit safely
}
// Branch B: Verifier updates R0 type to PTR_TO_MAP_VALUE
stat->packets += 1; // Safe!
When the verifier evaluates the conditional jump if R0 == 0, it branches its simulation state:
- On the
truepath:R0is known to be0. - On the
falsepath:R0is promoted toPTR_TO_MAP_VALUEwith known valid memory boundaries. Only here is dereferencing permitted.
4. JIT Compilation: From Bytecode to Native Machine Instructions
Once the verifier approves the bytecode, interpreting it instruction-by-instruction on every packet or syscall would introduce unacceptable CPU overhead.
The kernel runs the verified bytecode through an architecture-specific Just-In-Time (JIT) compiler (arch/x86/net/bpf_jit_comp.c on x86-64):
+--------------------+ eBPF JIT +--------------------+
| eBPF Bytecode | ─────────────────> | Native x86-64 |
| r1 = *(u64*)(r1+0) | | mov 0x0(%rdi),%rdi |
| r0 = 0 | | xor %eax,%eax |
| exit | | ret |
+--------------------+ +--------------------+
The JIT compiler maps eBPF virtual registers directly to physical hardware registers:
-
R0->%rax -
R1->%rdi -
R2->%rsi -
R3->%rdx -
R10->%rbp(stack frame pointer)
Once compiled into machine code:
- The memory page containing the generated instructions is allocated in kernel space.
- The kernel marks the page as read-only and executable (
PAGE_KERNEL_ROX). - Memory write permissions are permanently stripped from the page to prevent any privilege escalation or code-patching attacks.
5. Hook Attachment and Invocation
The compiled eBPF program exists in the kernel as an anonymous file descriptor (prog_fd). To run, it must attach to a kernel hook point.
Linux Kernel Space
┌────────────────────────────────────────────────────────────────────────┐
│ │
│ Network Device Driver ───> [ XDP Hook ] ───> Native eBPF Machine Code │
│ │ │
│ Syscall Entry (e.g. execve) ─> [ Tracepoint ] ─────┤ │
│ │ │
│ Kernel Function ─────────> [ Kprobe / Fentry ] ────┘ │
│ │
└────────────────────────────────────────────────────────────────────────┘
Common hook attachment targets include:
-
XDP (eXpress Data Path): Runs inside the network network interface card (NIC) driver before the Linux network stack even allocates an
sk_buff. -
Tracepoints: Static trace locations defined in kernel source code (for example,
sys_enter_execve). - Kprobes / Kretprobes: Dynamic insertion of breakpoints at arbitrary kernel function entries and returns.
- LSM (Linux Security Module): Security hooks capable of blocking unauthorized file, process, or network operations before they execute.
When the hook triggers (for example, a packet arrives at the NIC):
- The kernel prepares a context pointer (
ctx) containing packet data or CPU register states. - The kernel places this pointer into register
R1(%rdi). - It performs a direct hardware call instruction (
callq) to the JIT-compiled function address. - The eBPF function executes at full bare-metal CPU speed.
- The return code left in
R0(%rax) tells the kernel what to do next: pass the packet (XDP_PASS), drop it (XDP_DROP), or redirect it (XDP_TX).
6. Zero-Overhead Communication: BPF Maps and Ring Buffers
eBPF programs cannot call printf(), write directly to disk files, or open arbitrary TCP connections. So how do they communicate telemetry and state to user-space applications?
They use BPF Maps: kernel-managed, lockless memory structures shared between kernel and user space.
┌────────────────────────────────────────────────────────┐
│ User Space Application │
│ (Go / Rust / C Monitoring Daemon) │
└───────────────────────────▲────────────────────────────┘
│ mmap() / read()
│
┌───────────────────────────▼────────────────────────────┐
│ BPF Ring Buffer │
│ [Header Page] [Data Memory Page 1] [Data Page 2] │
└───────────────────────────▲────────────────────────────┘
│ bpf_ringbuf_submit()
│
┌───────────────────────────┴────────────────────────────┐
│ eBPF Kernel Program │
│ (Attached to tracepoint) │
└────────────────────────────────────────────────────────┘
The Old Approach: Perf Event Arrays
Historically, eBPF used BPF_MAP_TYPE_PERF_EVENT_ARRAY. This allocated a separate ring buffer for every individual CPU core.
- The Problem: Memory had to be over-provisioned across all cores to prevent drops on busy cores, while idle cores wasted RAM. If one core experienced a sudden spike, events were dropped even if total memory was mostly free.
The Modern Standard: BPF_MAP_TYPE_RINGBUF
Introduced in Linux 5.8, the BPF Ring Buffer provides a single, lockless, multi-producer single-consumer queue shared across all CPU cores:
- Kernel Side (Reservation and Submission):
struct event *e = bpf_ringbuf_reserve(&events, sizeof(*e), 0);
if (!e) {
return 0; // Drop if buffer is saturated
}
e->pid = bpf_get_current_pid_tgid() >> 32;
bpf_get_current_comm(&e->comm, sizeof(e->comm));
bpf_ringbuf_submit(e, 0);
-
User Space Side (Zero-Copy Polling):
The user-space daemon calls
mmap()on the ring buffer file descriptor. It monitors the buffer viaepoll()or a polling loop without issuing costlyread()system calls for individual events.
This architecture enables millions of kernel events per second to stream into user space with negligible CPU overhead and zero memory copies.
7. The Mental Model: eBPF in Summary
To build an accurate mental model of eBPF:
| Dimension | Loadable Kernel Module (LKM) | Traditional User-Space Agent | eBPF Program |
|---|---|---|---|
| Execution Location | Kernel space | User space | Kernel space |
| Safety Guarantee | None (Bug = Kernel Panic) | High (OS sandbox) | Absolute (Static verifier proof) |
| Execution Speed | Native CPU | Context-switch overhead | Native JIT-compiled CPU |
| Observability Depth | Complete access | Limited by syscalls | Raw tracepoints, XDP, and kprobes |
| Deployment Risk | Requires kernel recompilation/reboot | Easy to restart | Dynamic loading via bpf() syscall |
eBPF is not a virtual machine running on top of Linux. It is a compiler-verified, JIT-accelerated extension pipeline that allows you to safely inject programmable logic directly into the Linux kernel runtime.
Top comments (0)