Namespaces, capabilities, seccomp, and two 2026 kernel vulnerabilities that show where Docker isolation ends
Containers are lightweight because they do not run their own kernel.
They can have their own filesystem, process tree, network stack, and other isolated resources, but underneath all of that, containerized processes still use the host Linux kernel.
That one fact is extremely important for container security.
In 2026, two Linux kernel vulnerabilities made this boundary particularly interesting:
- CVE-2026-31431 ("Copy Fail") showed how a vulnerable kernel interface reachable from a container could affect resources outside that container.
-
CVE-2026-80521 demonstrated a container escape involving the Linux kernel's
AF_UNIXsubsystem.
This does not mean Docker is unsafe or that containers should not be used.
The important lesson is simpler:
Container isolation is ultimately enforced by the host kernel.
Understanding that boundary helps explain both the strengths and limitations of container security.
Containers Are Not Small Virtual Machines
A virtual machine runs its own guest operating system and its own kernel.
A container does not.
A container is better understood as a group of Linux processes running with restricted views and permissions provided by the kernel.
Those processes still make system calls to the host kernel.
This is one of the reasons containers are so lightweight. There is no separate guest kernel to boot for every container.
But it also creates an important security relationship:
The host Linux kernel is part of the trusted computing base for containers.
If a kernel vulnerability allows an attacker to bypass one of the mechanisms enforcing isolation, the container boundary can potentially be weakened or crossed.
Virtual machines are not completely immune to this problem either. Hypervisors have also had escape vulnerabilities.
The broader principle is that every isolation mechanism has a trusted computing base.
For conventional containers, a particularly important part of that trusted base is the Linux kernel.
What Actually Isolates a Container?
Docker uses several Linux security mechanisms. The most important ones include:
- Namespaces
- cgroups
- Capabilities
- seccomp
- AppArmor or SELinux
- User namespaces
Each solves a different problem.
Namespaces: What Can the Process See?
Linux namespaces provide processes with isolated views of system resources.
For example:
| Namespace | Main purpose |
|---|---|
| PID | Process visibility |
| Mount | Filesystem mount view |
| Network | Network interfaces and ports |
| UTS | Hostname |
| IPC | Inter-process communication |
| User | User and group ID mappings |
| Cgroup | Cgroup hierarchy view |
| Time | Certain process-visible clocks |
A simple way to remember namespaces is:
Namespaces control what a process can see.
A container can therefore have its own process list, filesystem view, and network environment even though those resources ultimately exist on the same kernel.
cgroups: How Much Can It Use?
Control groups, or cgroups, deal with resource management.
They can control or account for resources such as:
- CPU
- Memory
- Number of processes
- I/O
So the basic mental model is:
Namespaces
↓
What can this process see?
cgroups
↓
How much can this process use?
These mechanisms complement each other rather than solving the same problem.
Capabilities: What Privileged Operations Are Allowed?
Linux does not treat root as one giant permission.
Instead, many privileged operations are divided into capabilities.
Examples include:
CAP_NET_ADMIN
CAP_SYS_ADMIN
CAP_SYS_MODULE
CAP_SYS_PTRACE
Docker normally starts containers with a reduced set of capabilities.
That means a process can have UID 0 — effectively "root" inside the container — without automatically having every privileged operation available to host root.
Administrators can also explicitly add or remove capabilities:
--cap-add
--cap-drop
For application containers, dropping capabilities that are not required is generally preferable to giving the container additional privileges.
User Namespaces: Is Container Root Also Host Root?
User namespaces can map the container's UID 0 to an unprivileged UID on the host.
Conceptually:
Container
UID 0
↓
Host
UID 100000
Docker supports this through user namespace remapping, and rootless Docker provides another approach to reducing host-level privilege.
However, these protections depend on configuration.
On a normal Docker configuration without user namespace remapping, UID 0 inside the container can map directly to UID 0 on the host.
That does not mean the container automatically has unrestricted host access. Other controls such as capabilities, namespaces, seccomp, and LSM policies still apply.
But it does mean there is no additional UID-remapping boundary in that configuration.
seccomp: Which System Calls Can the Container Make?
A containerized application constantly interacts with the kernel through system calls.
seccomp can filter these calls.
Conceptually:
Container process
|
| system call
↓
seccomp filter
/ \
ALLOW DENY
|
↓
Linux kernel
Docker's default seccomp profile restricts access to a number of system calls.
This is useful because reducing the available system-call surface can make some kernel vulnerabilities harder to reach.
But there is an important limitation:
seccomp is not a complete sandbox.
It reduces the kernel interface available to a process. It does not make the remaining kernel code vulnerability-free.
That distinction becomes especially important when looking at real kernel vulnerabilities.
AppArmor and SELinux
Linux Security Modules, or LSMs, provide another layer of mandatory access control.
On systems using AppArmor, Docker can apply profiles such as docker-default.
Other distributions commonly use SELinux.
The overall security model therefore looks roughly like this:
Notice where the model ends:
the Linux kernel.
The kernel ultimately enforces these isolation mechanisms.
That is why kernel vulnerabilities can become container-security problems.
Case Study 1: Copy Fail — CVE-2026-31431
One useful example is CVE-2026-31431, known as "Copy Fail."
The vulnerability affected the Linux kernel's algif_aead component of the AF_ALG userspace cryptographic interface.
The important point for containers is not simply that the kernel contained a vulnerability.
It is that a default container could reach the vulnerable interface.
At a high level:
Container
↓
AF_ALG interface
↓
Linux kernel
↓
algif_aead
↓
CVE-2026-31431
According to Docker's technical analysis, the vulnerability could allow an unprivileged user with access to an AF_ALG socket to perform controlled writes to the page cache.
The page cache is a kernel-managed resource rather than something isolated per container.
That means the consequences are not necessarily limited to the container in which the vulnerable code executes.
The Important Concept: Reachability
This is one of the most useful ideas in container security.
A vulnerable kernel function is not automatically reachable from every container.
The practical question is:
Is vulnerable code present?
↓
Can the container reach it?
↓
Can the attacker trigger it?
↓
What resources can be affected?
For Copy Fail, the relevant interface was AF_ALG.
Docker therefore introduced additional security-profile restrictions to reduce container access to that interface.
But there is an important distinction:
A Docker security-profile mitigation does not repair vulnerable kernel code.
A kernel update fixes the underlying vulnerability.
A security-profile change can instead reduce the ability of containers to reach the vulnerable functionality.
Both approaches can be valuable.
Case Study 2: CVE-2026-80521 and AF_UNIX
A second example is CVE-2026-80521, involving the Linux kernel's AF_UNIX subsystem.
AF_UNIX sockets are widely used for local inter-process communication.
The published research described a use-after-free involving the kernel's AF_UNIX garbage-collection mechanism and demonstrated a container escape.
This example is interesting for a different reason.
AF_ALG is a relatively specialized interface.
AF_UNIX, on the other hand, is deeply integrated into normal Linux applications.
That creates a different defensive challenge.
You can potentially block access to a specialized interface.
It is much harder to remove access to a kernel subsystem that ordinary applications depend on.
A necessary distinction
It would be misleading to summarize this as:
"Containers can always be escaped."
That is far too broad.
A more accurate statement is:
A kernel vulnerability in functionality reachable from a container can potentially undermine the isolation mechanisms that container runtimes depend on.
Whether a particular system is exploitable depends on factors such as:
- Kernel version
- Whether the vulnerable code is present
- Whether the relevant interface is reachable
- Container configuration
- Capabilities
- seccomp
- LSM policies
- User namespace configuration
- Available mitigations
So the vulnerability demonstrates a limitation of shared-kernel isolation. It does not mean every container on every Linux system is automatically vulnerable.
What Does "Container Escape" Actually Mean?
A container escape occurs when code running inside a container gains access beyond the intended container boundary.
Depending on the vulnerability and configuration, that could mean access to:
- Host files
- Host processes
- Host kernel resources
- The container runtime
- Other workloads
Conceptually:
NORMAL
Container A ──X──> Host
Container A ──X──> Container B
AFTER A SUCCESSFUL ESCAPE
Container A ───────> Host
├── Host filesystem
├── Host processes
├── Container runtime
└── Other workloads
"Escape" is not one fixed outcome.
The actual impact depends heavily on the host configuration and what the attacker can access after crossing the intended boundary.
That is why saying simply "the container was escaped" does not tell the entire story.
A Small Hands-On Demonstration
I tested several of these concepts on my own CachyOS system using:
- CachyOS
- Linux kernel 7.2.7-1-cachyos
- Docker Engine 29.8.2
The following demonstrations do not exploit a vulnerability. They simply show how container isolation works.
Check Docker's Security Configuration
docker info --format "{{.SecurityOptions}} cgroup v{{.CgroupVersion}}"
My system reported:
[name=seccomp,profile=builtin name=cgroupns] cgroup v2
This showed that Docker was using its built-in seccomp profile and cgroup namespaces.
AppArmor was not enabled on this system.
That is an important reminder:
Container security depends partly on the host configuration.
Compare Container and Host Processes
Start a container:
docker run -d --rm --name nslab alpine sleep 30
Then inspect the process from the host:
ps -o user,pid,cmd -C sleep
The host can see the process.
Inside a normal container, however:
docker run --rm alpine ps
the container sees its own PID namespace rather than the host's complete process list.
This is a simple demonstration of what namespaces actually provide:
the process still exists on the host, but the container has a restricted view of it.
See What Happens With --privileged
Now compare a normal container with a privileged container:
docker run --rm --privileged alpine \
sh -c 'grep -E "^(CapEff|NoNewPrivs|Seccomp)" /proc/self/status'
A privileged container receives substantially broader privileges and disables the normal seccomp restrictions shown in the test.
This illustrates an important point:
"It runs in a container" does not tell you the security posture of that container.
You also need to know how it was configured.
Kubernetes: The Same Kernel Boundary Still Exists
Kubernetes does not create a second kernel for every Pod.
Pods ultimately run through a container runtime on a node and therefore share that node's Linux kernel.
However, Kubernetes security settings can significantly change the isolation provided to workloads.
For example, Kubernetes supports explicitly requesting the runtime's default seccomp profile:
spec:
securityContext:
seccompProfile:
type: RuntimeDefault
Other settings worth paying attention to include:
hostPID
hostNetwork
hostIPC
hostPath
privileged
These features can intentionally weaken parts of the normal isolation model.
So instead of asking:
"Does Kubernetes isolate my container?"
a better question is:
"Which isolation mechanisms are actually enabled for this workload and this node?"
Practical Hardening
The goal is not to make containers impossible to use.
The goal is to reduce unnecessary privileges and keep the underlying host secure.
Host
- Keep the host kernel patched.
- Reboot when an updated kernel needs to become active.
- Follow your distribution's security advisories.
- Stay on a supported kernel series.
Docker
- Keep Docker Engine updated.
- Avoid
--privilegedfor ordinary workloads. - Drop unnecessary capabilities.
- Run applications as non-root users where possible.
- Consider
no-new-privileges. - Keep the default seccomp profile enabled.
- Avoid
seccomp=unconfinedunless there is a documented reason. - Consider rootless Docker or user namespace remapping.
- Avoid mounting
/var/run/docker.sockunless it is genuinely required.
Kubernetes
- Use
seccompProfile: RuntimeDefault. - Consider enabling
seccompDefault. - Use
runAsNonRoot: true. - Set
allowPrivilegeEscalation: falsewhere appropriate. - Drop unnecessary capabilities.
- Avoid privileged containers.
- Avoid unnecessary host namespace sharing.
- Keep worker-node kernels patched.
What About Untrusted Workloads?
Not every workload needs the same isolation strength.
A normal web application running in a container has a different threat model from:
- Multi-tenant infrastructure
- Customer-submitted code
- Untrusted CI jobs
- Generated-code execution
- Security research sandboxes
For highly untrusted workloads, stronger isolation technologies may be appropriate.
Examples include:
- Firecracker
- Kata Containers
- gVisor
These approaches introduce their own trade-offs in complexity, compatibility, and performance.
The goal is not to declare containers obsolete.
The goal is to choose an isolation boundary appropriate for the workload.
What Defenders Can Actually Monitor
Patching is the first line of defense.
Detection is useful when prevention fails or when someone changes the security configuration unexpectedly.
On container hosts, useful signals can include unexpected:
setnsunshare- Mount operations
- Kernel module loading
- Access to container runtime sockets
- Privileged processes
- Changes to sensitive host files
In Kubernetes, defenders can also watch for unexpected:
privileged: truehostPIDhostNetworkhostIPChostPath- Capability additions
- Suspicious
kubectl exec - Changes to Pod security settings
Useful evidence sources include:
auditd- journald
- Container runtime logs
- Kubernetes audit logs
- Kernel logs
- eBPF-based runtime telemetry
Tools such as Falco and Tetragon can also provide runtime visibility into process and syscall activity.
For container escapes, node-level visibility is particularly important because the boundary being attacked is ultimately enforced by the node's kernel.
Conclusion
Containers remain one of the most useful ways to package and run software.
The important security lesson is not that containers are inherently unsafe.
It is that their isolation is implemented through kernel mechanisms.
The simplified model is:
Application
↓
Container configuration
↓
Capabilities / user identity
↓
seccomp / AppArmor / SELinux
↓
Linux kernel
↓
Hardware / virtualization boundary
The 2026 vulnerabilities discussed here illustrate two different sides of that boundary.
CVE-2026-31431 (Copy Fail) showed how restricting access to a vulnerable kernel interface can reduce container exposure.
CVE-2026-80521 demonstrated the harder case: when vulnerable functionality belongs to a subsystem that ordinary applications depend on, completely removing access to that functionality may not be practical.
The practical defense is therefore layered:
- Reduce container privileges.
- Restrict unnecessary kernel interfaces.
- Use seccomp and LSM policies.
- Keep the host kernel patched.
- Monitor the node, not only the application.
- Use stronger isolation when the workload is genuinely untrusted.
And the key idea to remember is simple:
A container can isolate an application from the host only as well as the kernel enforcing that isolation.



Top comments (0)