DEV Community

Ryo Tanaka
Ryo Tanaka

Posted on

How boxr runs rootless: namespaces, a trampoline, and uid mapping

How boxr runs rootless: namespaces, a trampoline, and uid mapping

boxr is a rootless OCI container engine in Rust. No daemon, no sudo, no setuid
helper. People often ask how the isolation actually works under the hood, so
here is the full story of what happens on Linux between you typing boxr run
and your process starting inside a container.

The short version: get into a user namespace as early as possible, then do
everything else from inside it.

The problem: threads and unshare do not mix

The boxr CLI is a multi-threaded tokio application. The kernel refuses
unshare(CLONE_NEWUSER) from any process that already has threads, it just
returns EINVAL. So the straightforward approach, calling unshare from the CLI
before starting containers, does not work. By the time your code runs, the
tokio runtime has already spawned its worker threads and the door is closed.

The trampoline

The fix is a re-exec. Before the tokio runtime starts, boxr re-execs itself
as __internal-trampoline. That subprocess is single-threaded, so it can
unshare the user namespace cleanly. All of the namespace setup described below
happens in the trampoline, in plain sequential code, before any async runtime
exists.

It is a small trick, but it sidesteps the entire class of problems that comes
from trying to do namespace setup inside an async runtime. Sequential code,
one thread, no surprises.

The user-namespace handshake

Getting the user namespace is a two-process dance:

  1. The trampoline forks. The child unshares CLONE_NEWUSER. At this point the child is root inside its own namespace, but it has no uid mapping yet, so it cannot do anything useful. Only the parent can write the maps.
  2. Parent and child synchronize over a Unix socketpair, a simple "ready" / "done" exchange, because the parent must write /proc/<pid>/uid_map for the child.
  3. The parent writes the uid/gid maps in RootlessUserConfig::setup_child_mappings. There are two paths:
    • If newuidmap and newgidmap exist and /etc/subuid plus /etc/subgid have ranges for the user, they are used. Container uid 0 maps to the host user, and container uids 1 and up map onto the subordinate range. This is the full multi-uid mapping.
    • Otherwise, boxr falls back to writing the maps directly: uid 0 maps to the host uid with a single entry, after writing "deny" to setgroups (the kernel requires that before it will accept a gid_map).
  4. The child is now root in its own user namespace, with real unprivileged uids everywhere outside it.

Network before the rest

With root-in-namespace secured, the child optionally unshares the network
namespace next. Networking is plumbed before the remaining namespaces on
purpose: the parent can attach a pasta process, or the pure-Rust usernet TAP
engine can spawn its worker, while the namespace layout is still simple.

The remaining namespaces and PID 1

Then the child unshares the PID, mount, UTS, and IPC namespaces. The cgroup
namespace is opt-in via annotation, and IPC and UTS can be left on the host
via annotations too. A second fork makes the grandchild PID 1 in the new PID
namespace. Finally:

  • bind-mount the rootfs onto itself (pivot_root requires it to be a mountpoint),
  • pivot_root into it (with a chroot fallback),
  • exec the container init.

The whole chain is: userns first, maps written by the parent, network plumbed,
the rest of the namespaces, pivot_root, exec. Nothing in it runs as root on
the host.

What is honest about the limits

  • Without subuid/subgid configured, the container sees a single mapped uid. Fine for dev containers, not a full multi-user mapping.
  • Rootless networking goes through pasta or the user-mode TAP engine, not a veth pair, so raw sockets and some packet types behave differently.
  • This is the native Linux path. The macOS and Windows runtimes take completely different routes (Virtualization.framework, WSL2).

If you work on container runtimes, the code is all in one place:
src/runtime/linux.rs has the trampoline and the namespace ordering, and
src/security/mod.rs has the uid/gid mapping strategy.

The repo is kchaitanya863/boxr.
Questions welcome, especially "why did you do it this way instead of X".

Top comments (0)