DEV Community

Cover image for What Is Copy-on-Write? How Can a Process "Copy" Memory Without Actually Copying It?
Aditya Sharma
Aditya Sharma

Posted on

What Is Copy-on-Write? How Can a Process "Copy" Memory Without Actually Copying It?

When a process calls fork(), the operating system creates a child process with what appears to be an identical copy of the parent's memory. If the parent had two gigabytes of working memory, did the kernel just allocate another two gigabytes?

The answer is no. And the reason why reveals something genuinely elegant about how operating systems manage memory.


What "Memory" Actually Means to a Process

Before getting to the trick, it helps to have a precise mental model of what a process's memory actually is.

A process does not own physical RAM the way a warehouse owns inventory. Instead, it operates through an abstraction: the virtual address space. Every memory address the process works with is a virtual address. When the CPU needs to read or write memory, it translates that virtual address to a physical address using a structure called a page table.

The page table is a mapping:

Virtual address
      ↓
Page table
      ↓
Physical address
Enter fullscreen mode Exit fullscreen mode

Memory is divided into fixed-size chunks called pages. A typical page is 4 KB. The page table maps virtual pages to physical pages (sometimes called frames). The CPU's memory management unit performs this translation automatically on every memory access.

The critical insight: two different processes can have virtual addresses that map to the same physical page. Their virtual address spaces are separate, but the underlying physical memory can be shared. This is not a bug or a security hole when the kernel controls it deliberately. It is the foundation that makes Copy-on-Write possible.


What fork() Actually Does to Memory

When a process calls fork(), the kernel creates a child with a logically equivalent virtual address space. The child sees the same variables, the same stack, the same heap, the same loaded libraries.

The naive approach: allocate new physical pages for everything, copy all the contents. For a process with gigabytes of memory, that means gigabytes of allocation and copying, taking real time and consuming real resources.

The Copy-on-Write approach: do not allocate new physical pages immediately. Instead, arrange for parent and child to reference the same physical pages through their respective page tables. Both virtual address spaces look identical, but they are temporarily pointing at the same underlying physical memory.

This works because the kernel controls the page-table mappings. The pages that both processes now share are marked with a protection that prevents writes through those mappings. If either process tries to modify a shared page, the CPU detects the permission violation and the kernel gets involved.

The processes have separate virtual address spaces. They just happen to have page-table entries pointing to the same physical pages, with those entries temporarily protected against modification.


The Moment It Matters: One Process Writes

The interesting part happens when a process actually tries to modify a shared page.

Suppose after fork(), the state looks like this:

BEFORE WRITE

Parent VA ──┐
            ├──→ Physical Page X
Child VA ───┘
Enter fullscreen mode Exit fullscreen mode

Both virtual-page mappings point to the same physical page. Both are marked read-only in the relevant page-table entries.

Now the child executes a write to that virtual address. The CPU checks the page-table permissions. The mapping says: not writable. The CPU raises a page fault.

Here is an important distinction: a page fault is not inherently an error. It is a hardware mechanism that tells the kernel the current memory access needs handling. The kernel examines the fault, determines it is a legitimate Copy-on-Write situation, and takes action.

The kernel's response, step by step:

  1. Allocate a new physical page.
  2. Copy the contents of the original shared page into the new page.
  3. Update the child's page-table mapping to point to the new physical page.
  4. Make the child's mapping writable.
  5. Leave the parent's mapping pointing to the original physical page, unchanged.
  6. Resume the faulting instruction. The child's write now completes against its own private copy. The parent is unaffected. The result:
AFTER CHILD WRITES

Parent VA ─────→ Physical Page X

Child VA ──────→ Physical Page Y
                  (copy of X)
Enter fullscreen mode Exit fullscreen mode

The CPU detected the protection violation and signaled the kernel. The kernel performed the allocation, copy, and table update. The whole sequence is transparent to both processes. From their perspective, they each have independent writable memory. From the kernel's perspective, most of the memory is still shared.


Why This Is Faster Than Copying Everything

Return to the original question. A process with two gigabytes of memory calls fork().

Without Copy-on-Write, the kernel would need to allocate two gigabytes of new physical pages and copy every byte before the child can run. The child then calls exec() to load a different program entirely, replacing its address space. The two-gigabyte copy was unnecessary.

With Copy-on-Write, fork() sets up page-table structures and marks shared pages appropriately. The physical memory is not duplicated at all yet. When the child then calls exec(), it replaces its address space with the new program. Most of the parent's physical pages were never copied because the child never wrote to them.

The combination of fork() and exec() is an extremely common Unix pattern. A shell spawns a child and immediately executes another program. Copy-on-Write makes this practical: the expensive duplication that seems implied by fork() mostly never happens.

When a process does actually modify memory after fork(), only the specific pages that were written get their own copies. If a two-gigabyte process modifies twenty megabytes worth of data, the system copies roughly twenty megabytes, not two gigabytes. The work is deferred until it is actually necessary, and then limited to exactly what became necessary.


Page Faults Are Not Just Errors

This is worth a moment on its own because it contradicts what many developers learn early.

A page fault means the CPU encountered a memory access that could not be completed by a straightforward lookup in the page table. That can happen for several reasons:

  • The page has not been loaded into physical memory yet (demand paging).
  • The page needs to be brought in from disk.
  • The access violates a permission, which might indicate Copy-on-Write.
  • The access truly is invalid, in which case the kernel delivers a segmentation fault to the process. The kernel inspects each page fault and decides what to do. A Copy-on-Write fault is handled transparently. A demand-paging fault loads the page and resumes execution. Only if the access is genuinely illegal does the process receive a signal.

Page faults are the operating system's hook for managing virtual memory lazily. They allow the kernel to defer work until the hardware tells it work is necessary. Copy-on-Write is one application of this mechanism. Demand paging is another. The page fault itself is just the hardware saying: "I need your help with this access."


What Happens If Both Processes Keep Writing

Copy-on-Write does not eliminate copying. It defers and minimizes it.

If a parent and child both write extensively to shared pages, both will eventually end up with their own independent physical pages. The initial sharing was useful because much of the memory might never have been touched. But write-heavy workloads can result in the full cost eventually being paid.

Copying also happens at page granularity. If a process modifies one byte within a page, the entire page typically needs to be copied before that write can proceed. This means Copy-on-Write can occasionally produce more copying than a naive byte-level approach if many single-byte writes land on many different shared pages. The performance benefit depends on how much sharing remains useful over time.

The kernel needs to know when a physical page is no longer shared. It keeps track of how many virtual-page mappings reference each physical page. When a page fault triggers a Copy-on-Write, the kernel checks this reference information. If only one process still references the page, that process can take ownership directly without allocating a new page, because there is nothing else to protect.


Copy-on-Write as a General Pattern

The same conceptual approach appears throughout systems design wherever deferred copying is useful.

Filesystem snapshots can record the current state of a volume without immediately duplicating all its data. Blocks are shared between the original filesystem and the snapshot until one of them modifies a block. Then that block gets its own copy. The snapshot reflects the old state. The live filesystem reflects the new state. Storage is consumed only as changes accumulate.

Virtual machine snapshots work similarly. The snapshot and the running VM can share memory pages and disk blocks until modification requires divergence.

Some container storage systems and copy-on-write filesystems use similar strategies for efficient image layering.

The common thread: start with shared immutable state, detect modification, diverge on demand. The implementations differ, and the mechanism for detecting modification differs, but the logical structure is the same as process Copy-on-Write.


Security and Isolation

Copy-on-Write relies on the kernel correctly enforcing the boundary between shared physical pages and independent writable mappings. The protection comes from:

  • page-table permission bits that prevent writes through shared mappings
  • CPU hardware that enforces those permissions on every access
  • kernel code that correctly handles Copy-on-Write faults and updates page tables accurately If the kernel handles a Copy-on-Write fault incorrectly, a process might write to physical memory that another process is still reading as shared. That would violate process isolation. The correctness of Copy-on-Write is therefore a correctness requirement for the kernel's virtual-memory system, not just a performance optimization.

This is why bugs in Copy-on-Write handling can be serious. A race condition or incorrect reference count in the Copy-on-Write path could potentially allow one process to corrupt another's memory. The physical page sharing that makes Copy-on-Write efficient also makes correctness essential.


Performance in Practice

The performance characteristics of Copy-on-Write are worth being honest about.

Benefits: fork() completes quickly even for large processes. Initial memory consumption after fork() is much lower than if everything were copied immediately. Workloads that fork() and exec() immediately pay almost no copying cost at all.

Costs: Each page fault during a Copy-on-Write operation takes real time: the CPU traps into the kernel, the kernel allocates a new page, copies the old page's contents, updates the page table, and resumes execution. For workloads that write to many shared pages, these page faults accumulate. In latency-sensitive contexts, the occasional pause for page copying can be noticeable.

Copy-on-Write is an optimization for common cases, not a guarantee that memory never needs to be copied. Its benefit is proportional to how much sharing remains useful after fork(). For processes that immediately diverge by modifying many pages, it buys less than for processes that mostly read shared state or quickly call exec().


The Mechanism, Assembled

The apparent paradox at the start of the article has an answer now.

fork() creates a child process whose virtual address space looks identical to the parent's. The kernel does not immediately allocate new physical pages for everything. It arranges for the parent's and child's page-table entries to reference the same physical pages, marking those entries as not writable through those mappings. The processes have separate virtual address spaces, but temporarily share physical memory.

When either process writes to a shared page, the CPU detects the permission violation and raises a page fault. The kernel examines the fault, determines it is a Copy-on-Write situation, allocates a new physical page, copies the original contents, updates the faulting process's page-table entry to point to the new page, marks it writable, and resumes execution. The two processes now have independent copies of that page.

The abstraction works because the virtual-memory system gives the kernel complete control over what each process sees when it accesses a given virtual address. Physical pages can be shared or separated, read-only or writable, present or absent, all without the process necessarily knowing. The hardware enforces the permissions and notifies the kernel when something needs to change.

Copy-on-Write works because the operating system does not ask "what if this memory changes?" It asks the hardware to tell it when the change actually happens, and then does exactly as much work as that particular change requires.

Top comments (0)