DEV Community

Cover image for What Actually Happens Inside OverlayFS: How Docker Layers, Copy-Up, and Whiteouts Really Work
Syed Anzar
Syed Anzar

Posted on

What Actually Happens Inside OverlayFS: How Docker Layers, Copy-Up, and Whiteouts Really Work

What Actually Happens Inside OverlayFS: How Docker Layers, Copy-Up, and Whiteouts Really Work

When you spin up a Docker container, it starts in milliseconds and appears to have its own complete, isolated filesystem. You can create files, modify system packages, and delete directories without altering the underlying host or the image from which the container was created.

Yet Docker is not creating a virtual hard disk or duplicating gigabytes of files on disk.

Underneath container runtimes like Docker, containerd, and Podman is OverlayFS, a Linux kernel union filesystem that merges multiple directory trees into a single, unified mount point.

While OverlayFS is remarkably efficient, treating it as a standard local filesystem leads to sudden latency spikes, unexplained disk inflation, and silent performance bottlenecks.

Here is what actually happens inside the Linux kernel when your containers read, write, copy, and delete files across layered filesystems.


1. The Core Architecture: The Four Overlay Directories

At its core, OverlayFS does not manage block storage. Instead, it sits on top of an existing native Linux filesystem (such as ext4 or xfs) and combines distinct directories into a single virtual hierarchy.

Every OverlayFS mount is defined by four core paths:

+-------------------------------------------------------------+
|                        merged/                              |
|   (The unified root filesystem visible inside the container)|
+-------------------------------------------------------------+
                              ▲
               +--------------+--------------+
               |                             |
      +-----------------+           +-----------------+
      |    upperdir/    |           |    lowerdir/    |
      |   (Read-Write)  |           |   (Read-Only)   |
      | Container layer |           | Image layers... |
      +-----------------+           +-----------------+
               ▲
               | (Atomic copy-up staging)
      +-----------------+
      |    workdir/     |
      |  (Kernel Temp)  |
      +-----------------+
Enter fullscreen mode Exit fullscreen mode
  1. lowerdir (Read-Only): A stacked list of base directories representing the container image layers. Lower directories are strictly immutable. A single OverlayFS mount can stack up to 128 lower directories (separated by colons).
  2. upperdir (Read-Write): The container's dedicated writable layer. Any new files created or modified by the running container are stored here.
  3. workdir (Internal Scratchpad): A private, empty directory on the same filesystem as upperdir. The Linux kernel uses this scratchpad to prepare files before moving them into upperdir, ensuring atomic operations across crashes.
  4. merged (The Unified Mount Point): The directory exposed to the container process. The kernel intercepts all VFS (Virtual File System) system calls in merged and routes them between upperdir and the lowerdir stack.

You can inspect this directly on any Linux host running Docker by checking active mounts:

mount | grep overlay
Enter fullscreen mode Exit fullscreen mode

The output reveals the exact parameters passed to the kernel:

overlay on /var/lib/docker/overlay2/<container-id>/merged type overlay (rw,relatime,
lowerdir=/var/lib/docker/overlay2/l/<id3>:/var/lib/docker/overlay2/l/<id2>:/var/lib/docker/overlay2/l/<id1>,
upperdir=/var/lib/docker/overlay2/<container-id>/diff,
workdir=/var/lib/docker/overlay2/<container-id>/work)
Enter fullscreen mode Exit fullscreen mode

2. Reading Files: The Lookup Chain and Page Cache Sharing

When a container process executes open("/etc/nginx/nginx.conf", O_RDONLY), the Linux VFS performs a top-down search:

  1. Check upperdir: The kernel checks whether /etc/nginx/nginx.conf exists in the container's writable layer.
  2. Scan lowerdir Stack: If absent from upperdir, the kernel walks through each layer specified in lowerdir from newest (top) to oldest (bottom) until it finds a matching entry.
  3. Cache the Dentry: The kernel caches the directory entry (dentry) pointing directly to the underlying physical inode on disk.
Read Request: open("/app/config.json", O_RDONLY)
  │
  ├─► Check upperdir/app/config.json ──────► Found? ──► Return fd (Container modified)
  │                                           │ (No)
  └─► Walk lowerdir stack (top to bottom) ────┘
        ├─► Layer 3: not found
        └─► Layer 2: found! ────────► Open lower file directly (Shared Page Cache)
Enter fullscreen mode Exit fullscreen mode

The Performance Advantage: Zero-Copy Page Cache

Because files in lowerdir are opened directly by the kernel without copying, multiple running containers sharing the same base image read the exact same physical pages in RAM.

If you run 50 microservice containers on a single host using the same Python runtime or Node.js binary, the Linux kernel loads those binary pages into the page cache only once.


3. Modifying Files: The Hidden Penalty of "Copy-Up"

What happens when your application opens an existing file from a lower layer in write mode?

open("/var/log/app.log", O_WRONLY | O_APPEND);
Enter fullscreen mode Exit fullscreen mode

Because lowerdir is strictly read-only, the kernel cannot write changes to the original file. Instead, OverlayFS intercepts the system call and triggers an internal kernel routine called ovl_copy_up().

            Step 1: Create temp inode
            ┌───────────────────────┐
            │  workdir/temp_file    │
            └───────────┬───────────┘
                        │
                        ▼ Step 2: Synchronous data + xattr copy
┌───────────────────────┐           ┌───────────────────────┐
│ lowerdir/database.db  ├──────────►│  workdir/temp_file    │
│ (2GB base layer file) │           └───────────┬───────────┘
└───────────────────────┘                       │
                                                ▼ Step 3: fsync() & atomic rename
                                    ┌───────────────────────┐
                                    │  upperdir/database.db │
                                    └───────────────────────┘
Enter fullscreen mode Exit fullscreen mode

The copy-up process follows strict atomic stages:

  1. Allocate in workdir: The kernel creates a temporary inode in the workdir directory.
  2. Synchronous Data Copy: The kernel copies the entire file content, file ownership, permissions, timestamps, and extended attributes (xattrs) from lowerdir into the temporary file in workdir.
  3. Flush to Disk (fsync): OverlayFS invokes fsync(2) on the temporary file to ensure that all copied data and metadata are durably written to physical media.
  4. Atomic Rename: The kernel atomically renames the temporary file from workdir to its final path in upperdir.
  5. Resume Syscall: The kernel switches the file descriptor to point to the new inode in upperdir and allows the original write syscall to complete.

Why Copy-Up Causes Unexpected Latency

Copy-up operates at the file level, not at the disk block level.

If your base image contains a 4GB pre-seeded database file or machine learning model, and a container process writes a single byte to it:

  • The process blocks synchronously while the kernel reads 4GB from disk and writes 4GB into workdir.
  • The process blocks while fsync() flushes 4GB to storage.
  • Only then does the 1-byte write execute.

If this happens during application startup or in a high-throughput request loop, you will see sudden thread starvation and severe I/O latency spikes.


4. Deleting Files: The Whiteout Mechanism

If lower layers are read-only, how can a container delete an inherited file with rm /app/secret.txt?

The kernel cannot remove the file from lowerdir. Instead, OverlayFS creates a Whiteout Device in upperdir.

A whiteout is a special character device node created with major number 0 and minor number 0 (makedev(0, 0)):

# Inside the upperdir on the host:
ls -l /var/lib/docker/overlay2/<id>/diff/app/
Enter fullscreen mode Exit fullscreen mode
c--------- 1 root root 0, 0 Sep 30 18:00 secret.txt
Enter fullscreen mode Exit fullscreen mode

When the container lists or opens /app/secret.txt:

  1. The kernel checks upperdir and sees the character device 0/0 named secret.txt.
  2. OverlayFS recognizes this device as a whiteout marker.
  3. The kernel suppresses the file: it conceals the lower file, hides the whiteout device itself from userspace directory listings (readdir), and returns ENOENT (No such file or directory) to any read request.
Container View (/merged)           Underlying Host Storage
┌──────────────────────┐           ┌────────────────────────────────────────┐
│                      │           │ upperdir/                              │
│  app/                │           │   └── app/                             │
│   (secret.txt is     │  ◄──────  │        └── secret.txt (char dev 0/0)   │
│    invisible)        │           │                                        │
│                      │           │ lowerdir/                              │
│                      │           │   └── app/                             │
│                      │           │        └── secret.txt (100MB data)     │
└──────────────────────┘           └────────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

The Dockerfile Anti-Pattern: Why rm -rf Doesn't Shrink Images

This whiteout mechanism explains a classic developer blunder in multi-line Dockerfiles:

# ANTI-PATTERN
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y build-essential curl  # Layer 1 (+400MB)
RUN rm -rf /var/lib/apt/lists/*                                # Layer 2 (+0MB saved!)
Enter fullscreen mode Exit fullscreen mode

In Layer 1, the package lists are committed into lowerdir.

In Layer 2, rm -rf creates hundreds of 0-byte whiteout device files in upperdir. The original 400MB of package lists remains fully intact in Layer 1. The resulting image size does not decrease by a single byte.

To actually reduce image size, deletion must occur in the same RUN step:

# CORRECT
FROM ubuntu:24.04
RUN apt-get update && apt-get install -y --no-install-recommends \
    build-essential \
    curl \
    && rm -rf /var/lib/apt/lists/*
Enter fullscreen mode Exit fullscreen mode

5. Deleting & Replacing Directories: Opaque Attributes

What happens when a container deletes an entire directory inherited from an image and creates a new directory with the same name?

If you run:

rm -rf /opt/config && mkdir /opt/config
Enter fullscreen mode Exit fullscreen mode

Creating a whiteout device for every single past and future file inside /opt/config would be slow and fragile. Instead, OverlayFS uses extended filesystem attributes (xattrs).

When an upper directory replaces a lower directory, the kernel sets an extended attribute on the new directory in upperdir:

trusted.overlay.opaque = "y"
Enter fullscreen mode Exit fullscreen mode

(In unprivileged rootless containers, this is stored under user.overlay.opaque="y".)

When the VFS walks the directory tree:

  • If a directory in upperdir has opaque = "y", the kernel stops the lookup immediately.
  • It completely ignores any directory with the same name in all lower layers, preventing lower files from leaking into the new directory.

6. The chmod -R Trap: Multiplying Container Disk Space

One of the most expensive operations in containerized workflows is bulk permission modifications.

Consider this common Dockerfile instruction or entrypoint script:

COPY ./app /app
RUN chown -R node:node /app && chmod -R 755 /app
Enter fullscreen mode Exit fullscreen mode

Because file ownership and permission bits are stored in the inode metadata, changing permissions on an existing lower-layer file requires modifying its inode.

Before modern kernel optimizations, changing the permissions of 10,000 files in /app forced the kernel to execute ovl_copy_up() on all 10,000 files:

  • 10,000 full file reads from lowerdir
  • 10,000 full file writes to workdir
  • 10,000 synchronous fsync() calls
  • 10,000 atomic renames into upperdir

If /app contained 500MB of dependencies (like node_modules), the image gained an unnecessary 500MB duplicate layer.

The Modern Mitigations

  1. COPY --chown: Always set ownership during copy rather than in a separate RUN step:
   COPY --chown=node:node ./app /app
Enter fullscreen mode Exit fullscreen mode
  1. OverlayFS metacopy (Linux Kernel 4.19+): Modern Linux kernels support metadata-only copy-up (metacopy=on). When a process only modifies permissions or timestamps, the kernel creates a lightweight stub in upperdir with the attribute trusted.overlay.metacopy instead of copying the file's data payload. Data is only copied if the file content is subsequently modified.

7. Direct Comparison: OverlayFS vs Native Filesystem

Operation Native Ext4 / XFS OverlayFS Container Layer
Read Unmodified File Direct inode lookup Walks upperdir → lowerdir stack (dentry cached)
Write New File Standard inode allocation Direct write to upperdir
First Write to Existing File In-place block write / append Full synchronous ovl_copy_up via workdir + fsync
Delete Inherited File Inode unlinked, blocks freed Writes c 0 0 whiteout device; lower data remains
RAM Cache for Shared Files Per-file page cache Single shared page cache across all containers
Inode Stability (st_dev) Inode number stays constant Inode / device ID can shift during copy-up without xino=on

Summary & Architectural Rules

Understanding OverlayFS turns container storage from a black box into a predictable system:

  1. Never write heavy runtime data to container layers: Database storage engines, message queues, and large log writers should always use Docker Volumes or Bind Mounts. Volumes bypass OverlayFS entirely, writing directly to native host filesystems without copy-up latency.
  2. Combine build and cleanup operations in single Dockerfile RUN steps: Deleting files in a separate layer only writes 0-byte whiteouts without reclaiming lower-layer space.
  3. Avoid mass permission rewrites (chmod -R / chown -R): Use COPY --chown to prevent layer duplication.
  4. Take advantage of read sharing: Identical read-only lower layers share host memory page caches across dozens of running containers.

Verified References & Kernel Sources

Top comments (0)