DEV Community

Cover image for How Linux Zero-Copy Actually Works: sendfile, DMA Scatter-Gather, and Page Cache Internals
Syed Anzar
Syed Anzar

Posted on

How Linux Zero-Copy Actually Works: sendfile, DMA Scatter-Gather, and Page Cache Internals

When building high-throughput web servers, reverse proxies, or message brokers like Kafka and NGINX, you inevitably run into the term Zero-Copy.

Most developers understand the high-level pitch: Zero-copy transfers data from disk to network without wasting CPU cycles copying bytes into user space.

While that summary is accurate, it skips the actual engineering underneath:

  • Where do the copies actually happen in a traditional read() and write() loop?
  • How does the kernel send data it never copied into memory?
  • Why did HTTPS break zero-copy for nearly a decade, and how did the Linux kernel fix it?

Let's trace what happens inside the Linux kernel, CPU registers, page cache, and network interface card (NIC) when you move bytes with and without zero-copy.


1. The Hidden Tax of Traditional I/O (read + write)

Consider the standard way a beginner writes a static file server in C, Python, or Go:

char buffer[8192];
while ((bytes_read = read(file_fd, buffer, sizeof(buffer))) > 0) {
    write(socket_fd, buffer, bytes_read);
}
Enter fullscreen mode Exit fullscreen mode

To transfer one chunk of data, the OS goes through 4 context switches and 4 data copies:

+-----------------------------------------------------------------------+
| USER SPACE                                                            |
|                           [ User Buffer ]                             |
|                             ^         |                               |
|                     CPU Copy|         |CPU Copy                       |
|                             |         v                               |
+-----------------------------|---------|-------------------------------+
| KERNEL SPACE                |         |                               |
|                     +-------+         +-------+                       |
|                     |                         |                       |
|                     v                         v                       |
|              [ Page Cache ]             [ Socket Buffer ]             |
|                     ^                         |                       |
|              DMA Copy|                         |DMA Copy               |
|                     |                         v                       |
|                  [ Disk ]                   [ NIC ]                   |
+-----------------------------------------------------------------------+
Enter fullscreen mode Exit fullscreen mode

The Step-by-Step Breakdown

  1. read() Syscall (User -> Kernel Switch 1): The CPU traps into kernel mode. The kernel issues a DMA (Direct Memory Access) request to read the file from storage into kernel page cache memory (Copy 1: Disk -> Page Cache).
  2. Kernel -> User Copy (Switch 2): The CPU copies the raw bytes from the kernel page cache into the process's stack or heap buffer (Copy 2: Page Cache -> User Buffer). The syscall returns to userspace.
  3. write() Syscall (User -> Kernel Switch 3): The process calls write(), trapping back into kernel mode. The CPU copies data from the user buffer into the kernel's socket send buffer (Copy 3: User Buffer -> Socket Buffer).
  4. Socket -> NIC DMA (Switch 4): The kernel returns from write(). Asynchronously, the NIC engine reads the socket buffer via DMA and sends the packets across the physical wire (Copy 4: Socket Buffer -> NIC).

Why Is This Expensive?

The DMA copies (Disk to RAM, RAM to NIC) are handled by hardware controllers without consuming CPU cycles.

The real bottleneck is the 2 intermediate CPU copies across the kernel-userspace boundary. When serving gigabytes of data per second:

  1. The CPU spends billions of clock cycles copying memory blocks (memcpy).
  2. It saturates the system memory bus.
  3. It evicts valuable hot data from the CPU L1, L2, and L3 caches to make room for transient network payloads.

2. Evolution Level 1: mmap() + write()

The first optimization historically attempted was memory-mapped I/O via mmap():

void *file_ptr = mmap(NULL, file_size, PROT_READ, MAP_SHARED, file_fd, 0);
write(socket_fd, file_ptr, file_size);
Enter fullscreen mode Exit fullscreen mode

By mapping the kernel's page cache directly into the process's virtual address space, mmap() eliminates the copy into the user buffer:

  • Disk -> Page Cache (DMA)
  • Page Cache -> Socket Buffer (CPU Copy)
  • Socket Buffer -> NIC (DMA)

This cuts down CPU copies from 2 to 1.

The Hidden Trap: SIGBUS Crashes

If another process truncates or deletes the file while your thread is executing write(socket_fd, file_ptr, ...), accessing that memory address generates a SIGBUS (Bus Error) signal. Unless your application installs a custom signal handler, the process crashes instantly.


3. Evolution Level 2: sendfile() Without Hardware Support

Linux 2.2 introduced the sendfile() syscall:

#include <sys/sendfile.h>

ssize_t sendfile(int out_fd, int in_fd, off_t *offset, size_t count);
Enter fullscreen mode Exit fullscreen mode

sendfile() instructs the kernel to transfer count bytes directly from in_fd (a file) to out_fd (a network socket).

+-----------------------------------------------------------------------+
| USER SPACE                                                            |
|                      (No data enters userspace)                       |
+-----------------------------------------------------------------------+
| KERNEL SPACE                                                          |
|              [ Page Cache ] ---- CPU Copy ----> [ Socket Buffer ]     |
|                     ^                                  |              |
|              DMA Copy|                                  |DMA Copy      |
|                     |                                  v              |
|                  [ Disk ]                            [ NIC ]          |
+-----------------------------------------------------------------------+
Enter fullscreen mode Exit fullscreen mode

Results:

  • Context switches dropped from 4 to 2 (only one syscall invocation).
  • Zero user-space memory allocations.
  • But on older kernels or basic hardware, there is still 1 CPU copy between the page cache and the socket buffer.

4. Evolution Level 3: True Zero-Copy (Scatter-Gather DMA)

Starting in Linux 2.4, the kernel introduced Scatter-Gather DMA support for network sockets (NETIF_F_SG).

When the network card supports scatter-gather, the CPU no longer copies file data into the socket buffer. Instead, the kernel only passes buffer descriptors (pointers and lengths) to the socket buffer:

+-----------------------------------------------------------------------+
| USER SPACE                                                            |
|                      (No data enters userspace)                       |
+-----------------------------------------------------------------------+
| KERNEL SPACE                                                          |
|              [ Page Cache (Physical RAM) ]                            |
|                     ^           |                                     |
|              DMA Copy           | Direct DMA Read                     |
|              from Disk          |                                     |
|                     |           v                                     |
|                  [ Disk ]    [ NIC ] <--- Packet Headers              |
|                                 ^                                     |
|                                 | DMA Read                            |
|                       [ Socket Buffer ]                               |
|                 (File Descriptors & Headers Only)                     |
+-----------------------------------------------------------------------+
Enter fullscreen mode Exit fullscreen mode

How struct sk_buff Handles This Internally

Inside the kernel networking stack, every packet is represented by a struct sk_buff (socket buffer).

Instead of allocating contiguous memory to hold both the TCP/IP headers and the file payload, the kernel uses fragmented skb buffers (skb_shinfo):

struct skb_shared_info {
    /* ... */
    nr_frags;
    skb_frag_t frags[MAX_SKB_FRAGS];
};
Enter fullscreen mode Exit fullscreen mode

Each skb_frag_t holds:

  • A pointer to the physical memory page (struct page *) in the Page Cache.
  • An offset inside that page.
  • The byte length.

When the packet is scheduled for transmission:

  1. The kernel constructs the TCP/IP headers in memory.
  2. The NIC's DMA engine receives a scatter-gather list:
    • Descriptor 1: Read 54 bytes (TCP/IP headers) from kernel memory.
    • Descriptor 2: Read 1460 bytes (file payload) directly from the page cache address.
  3. The NIC controller gathers the memory regions over the PCIe bus and stitches them into a single Ethernet frame on the wire.

CPU copies performed: Exactly ZERO.


5. splice(): Zero-Copy Between Arbitrary File Descriptors

While sendfile() is designed specifically for streaming files to sockets, Linux 2.6.17 introduced splice() to support zero-copy streaming between arbitrary file descriptors (pipes, sockets, block devices):

#define _GNU_SOURCE
#include <fcntl.h>

ssize_t splice(int fd_in, loff_t *off_in, int fd_out, loff_t *off_out, 
               size_t len, unsigned int flags);
Enter fullscreen mode Exit fullscreen mode

The Magic of Pipe Buffers (struct pipe_buffer)

splice() works using kernel pipe buffers (struct pipe_inode_info). A pipe buffer in Linux does not allocate memory for data bytes; it is a circular ring of page pointers:

struct pipe_buffer {
    struct page *page;
    unsigned int offset;
    unsigned int len;
    const struct pipe_buf_operations *ops;
    unsigned int flags;
};
Enter fullscreen mode Exit fullscreen mode

When you splice from a socket into a pipe, the kernel pins the network incoming page into the ring buffer. When you splice from that pipe into a file or another socket, the kernel transfers ownership of those page pointers without copying a single byte.


6. The Modern HTTPS Problem: Why TLS Broke Zero-Copy

For years, zero-copy had a massive limitation: Encryption.

If your web server runs HTTPS, traditional OpenSSL or BoringSSL pipelines look like this:

  1. Read encrypted/raw data into userspace.
  2. Run AES-GCM encryption in user memory.
  3. Call write() with the ciphertext.

Because encryption requires reading and modifying every single byte, sendfile() was useless for HTTPS. Web servers were forced back into the slow 4-switch, 2-copy path.

The Solution: kTLS (Kernel TLS)

Linux 4.13 introduced kTLS (Kernel TLS, via SO_TLS_TX / TCP_ULP):

/* Enable Kernel TLS User-Level Protocol on the TCP socket */
setsockopt(sock_fd, SOL_TCP, TCP_ULP, "tls", sizeof("tls"));

/* Provide the symmetric cipher keys (AES-GCM-128 / 256) negotiated in userspace */
setsockopt(sock_fd, SOL_TLS, TLS_TX, &crypto_info, sizeof(crypto_info));
Enter fullscreen mode Exit fullscreen mode

Once kTLS is enabled:

  1. The userspace library (such as OpenSSL or Rustls) completes the initial TLS 1.3 handshake.
  2. The session encryption keys are handed off to the kernel socket options.
  3. The application calls standard sendfile()!
  4. The kernel encrypts the page cache buffers inline as they flow through the network stack, or offloads encryption entirely to SmartNICs with TLS hardware accelerators.

This restored zero-copy throughput for HTTPS static file delivery in modern NGINX, HAProxy, and Envoy deployments.


7. Zero-Copy Performance Summary

Mechanism Syscalls Used Context Switches CPU Memory Copies DMA Copies TLS Compatible?
Traditional I/O read() + write() 4 2 2 Yes (Userspace crypto)
Memory Mapped mmap() + write() 4 1 2 Yes (Userspace crypto)
sendfile() sendfile() 2 1 (0 with SG-DMA) 2 Only with kTLS
splice() splice() 2 0 2 Only with kTLS
MSG_ZEROCOPY send(..., MSG_ZEROCOPY) 2 0 (Page pinning) 1 App-dependent

8. Senior Engineer Gotchas to Keep in Mind

  1. MSG_ZEROCOPY on small payloads is a trap: For payloads under 10-16 KB, standard send() is faster. MSG_ZEROCOPY on raw memory buffers requires the kernel to pin user pages (get_user_pages_fast), set up completion notifications on the socket error queue (sock_extended_err), and notify userspace via poll()/epoll(). The bookkeeping overhead easily exceeds the cost of a small memcpy.
  2. File modification during transfer: If you modify a file in place while sendfile() is actively reading page cache pages, the client will receive corrupted, mixed data frames. Always use atomic file replacements (rename/symlinks) for live assets.
  3. Dirty Page Flushing: Zero-copy does not bypass Linux caching rules. Data must exist in the kernel page cache; if pages are evicted under memory pressure, synchronous disk reads will stall the I/O thread.

Key Takeaway

Zero-copy is not about avoiding disk or network hardware transfers; it is about eliminating redundant CPU memory copies and context switches across the kernel-userspace boundary.

By passing page descriptors (struct page *) through struct sk_buff and leveraging NIC Scatter-Gather DMA, the Linux kernel lets hardware controllers stream data straight from disk cache to the wire while keeping your CPU caches clean and cool.

Top comments (0)