A Linux process can be compromised without giving the attacker unlimited access to the operating system.
That statement might seem optimistic, but it points to a real mechanism. Most people think of Linux security in terms of permissions: who owns the file, what group a process runs as, whether a user has root. Those things matter. But there is another dimension of control that is less visible and, in many ways, more fundamental: restricting what a process is allowed to ask the kernel to do.
Even a compromised process has to work through the same kernel interface as a legitimate one. If you can control which requests that interface accepts, you can limit the damage an attacker can cause.
That is what seccomp does.
System Calls: The Kernel Interface
Before explaining seccomp, it helps to be precise about system calls, because seccomp works at exactly that boundary.
Every user-space process runs in a restricted execution environment. It can execute instructions, access memory it owns, and perform arithmetic. But it cannot directly perform privileged operations: reading a file, creating a socket, spawning a process, or communicating with hardware. To do any of those things, the process has to ask the kernel.
A system call is that request. The process prepares a request, identifies which kernel operation it needs by number, and transitions execution into kernel space. The kernel validates the request, carries it out, and returns a result. Then execution returns to user space.
User Space
---------------------------------
Application
↓
libc / syscall interface
↓
system call (e.g., read, socket, fork)
---------------------------------
Kernel Space
↓
Kernel validates and handles request
↓
Result returned to process
---------------------------------
This is the mechanism behind almost everything a process does. Open a file: system call. Send a network packet: system call. Allocate memory from the OS: system call. The kernel is the gateway to everything outside the process's own address space.
Why System Calls Are a Security Boundary
A typical Linux system exposes hundreds of distinct system calls. A process running on that system has access to all of them by default, subject to the usual permission checks.
That creates a large interface. Some of those calls are mundane. Others can mount filesystems, load kernel modules, reboot the machine, trace arbitrary processes, or modify kernel behavior. A process that needs to read and write files doesn't need the calls that do any of those things.
The security principle here is straightforward: the more kernel functionality a process can reach, the more an attacker who controls that process can potentially do. A process that only needs five system calls to do its job has no reason to have access to five hundred.
This is the motivation behind seccomp. Not "what user is this process?" or "what files can it access?" but rather: "which kernel operations is this process even allowed to request?"
What Seccomp Does
Seccomp (Secure Computing) is a Linux kernel mechanism that allows a process to restrict the system calls it can make. Once applied, a seccomp policy sits between the process and the kernel. When the process attempts a system call, the kernel checks the policy before deciding what to do.
Process
↓
"Attempt system call X"
↓
Kernel boundary
↓
Seccomp policy check
↓
┌─────────────────────┐
│ Allowed? → continue│
│ Denied? → action │
└─────────────────────┘
Seccomp does not intercept arbitrary user-space instructions. A compromised process can still execute code in its own address space, modify its own data, or perform computations. What seccomp restricts is specifically the transition into kernel space through system calls. That boundary is where the process interacts with the rest of the system.
Strict Mode and Filter Mode
Seccomp has two operational modes, and they represent a significant evolutionary step.
The original strict mode is extremely limiting. A process in strict mode can only invoke a handful of system calls: read, write, exit, and signal-return. Anything else causes the process to be killed. This is useful in very narrow cases where an application genuinely needs only those operations after a certain point.
Filter mode is the important model for modern use. Instead of a fixed allowlist, filter mode lets a process (or a parent process acting on its behalf) install a custom policy describing which system calls are permitted, which are rejected, and what happens in each case.
Filter mode is what makes seccomp practical for real applications.
Seccomp-BPF: Flexible Filtering
Filter mode uses a BPF-based mechanism to evaluate system call requests. BPF (Berkeley Packet Filter) was originally designed for network packet filtering, but its instruction set was adapted for seccomp because it provides a way to express filtering logic that the kernel can evaluate efficiently.
It is worth being precise: seccomp does not run arbitrary eBPF programs with full kernel access. Seccomp-BPF uses a restricted instruction set applied specifically to system-call metadata. The filter can inspect the system call number, the architecture identifier, and certain system call arguments. It cannot perform I/O, access kernel data structures freely, or take arbitrary actions. The scope is intentionally narrow.
Process
↓
system call
↓
Kernel
↓
seccomp-BPF filter evaluates:
- syscall number
- architecture
- argument values (where applicable)
↓
Filter returns a decision
├── Allow
├── Return error to process
├── Trap (generate signal)
├── Kill thread or process
└── Notify user space (where supported)
↓
Kernel acts on decision
The filter is evaluated every time the process attempts a system call. If the call is on the policy's allowed list, execution continues normally. If not, the kernel takes whatever action the policy specified.
What a Filter Can Inspect
A seccomp filter has access to information associated with the system call attempt:
The system call number. Linux identifies each system call by a unique number. read is one number, write another, socket another. The filter can allow or deny based on which call the process is trying to make.
The architecture identifier. This is relevant because system call numbers are architecture-specific. The same number can mean different operations on different CPU architectures. Filters commonly check the architecture to prevent mismatches between what the filter expects and what the hardware might interpret.
System call arguments. Some filtering policies care not just about which call is being made, but how it is being made. A filter might permit socket but only for certain address families, or permit ioctl only with specific request codes. Argument filtering requires care because of how certain argument types are passed, and not every argument can be safely filtered in every case.
These three inputs give a filter meaningful control over the system call interface without needing to understand the full semantics of every possible operation.
Seccomp Actions
When a filter matches a system call, it can specify different outcomes:
Allow (ALLOW). The system call proceeds. The kernel handles it as it normally would.
Return error (ERRNO). The system call fails and returns a specified error code to the process. The process sees a failure but continues running. This is useful when you want to deny an operation without killing the process.
Trap (TRAP). The kernel delivers a signal to the process. This is typically used when a monitoring or supervisor process wants to intercept the attempt and handle it in user space.
Kill thread (KILL_THREAD). The specific thread that made the system call is immediately killed.
Kill process (KILL_PROCESS). The entire process is terminated. This is appropriate when an unexpected system call indicates the process has been compromised and should not continue.
Notify (USER_NOTIF). In supported configurations, the syscall attempt is forwarded to a supervisor process in user space, which can make the final decision. This enables more sophisticated policy designs.
The choice of action reflects how the policy designer wants to handle violations. A process that is expected to be well-behaved might use ERRNO for gradual fail-safe behavior. A process handling sensitive operations in a high-security context might use KILL_PROCESS to ensure that any unexpected system call results in immediate termination.
A Concrete Example: Sandboxed PDF Parser
Consider a PDF parsing service. Its job is to accept a PDF file, read its contents, parse the document structure, and produce output. It processes untrusted input, which means it is a prime target for exploitation if a malicious PDF triggers a parser bug.
A seccomp policy for this process might look conceptually like this:
PDF Parser
↓
Required operations:
├── read
├── write
├── mmap (memory management)
├── close
└── exit
Everything else: return error or kill
If an attacker exploits a bug in the parser and tries to use the compromised process to create a socket, fork a child process, load a kernel module, or reboot the machine, the seccomp filter intercepts those system calls before they reach the kernel. The attacker controls the process but cannot use it to reach most of the system.
This does not make the parser invulnerable. An attacker might still exfiltrate data through allowed write operations or find another way to achieve their goal within the permitted syscall set. But the attack surface is dramatically reduced.
Seccomp Is One Layer, Not a Complete Sandbox
It is important to be explicit about what seccomp does not do.
Seccomp restricts the system calls a process can make. It does not:
- Restrict what the process can read or write in its own memory
- Control what files or network resources the process can access through allowed syscalls
- Replace filesystem permissions
- Prevent all possible exploitation
- Protect against kernel vulnerabilities that are reachable through allowed calls Modern Linux isolation uses multiple independent mechanisms together:
Process
|
┌─────────┼──────────┐
↓ ↓ ↓
Seccomp Namespaces Capabilities
↓ ↓ ↓
Syscalls Visibility Privileges
|
Combined isolation
Namespaces control what a process can see: its own view of filesystems, networks, process IDs, users. They create the appearance of isolation by restricting visibility.
Capabilities determine which privileged operations a process can perform, such as opening privileged network ports, loading kernel modules, or changing system time.
Seccomp restricts which kernel interfaces the process can invoke at all.
These mechanisms operate at different layers. A process might have a capability to perform a privileged operation but lack seccomp permission to even invoke the syscall that would use it. A process might be in an isolated namespace but still capable of making system calls that affect the host kernel unless seccomp is also applied.
Effective sandboxing typically combines all of them.
Seccomp in Containers
Containers are not virtual machines. A container runs using the host kernel, sharing the same underlying system as everything else on the machine. Isolation comes from namespaces, cgroups, capabilities, and similar mechanisms, not from a separate kernel.
This makes seccomp particularly relevant. Every containerized process is making system calls to the same kernel as the host system. A containerized process with access to dangerous system calls can potentially affect the kernel, regardless of what namespace it lives in.
Container runtimes often apply seccomp profiles to reduce the syscall surface available to containers. The default profiles for tools like Docker or containerd allow a broad but pruned set of system calls: enough for most applications, but with some of the more dangerous kernel interfaces removed.
Applications with specific needs can use more restrictive custom profiles. The more precisely a profile matches what an application actually requires, the smaller its exposed kernel interface.
Seccomp and Browser Sandboxing
Browsers handle untrusted content from across the internet. A vulnerability in a browser's rendering engine, if exploited, gives an attacker control of the rendering process. The question is how much damage that control allows.
Modern browsers often isolate rendering in separate processes with restricted kernel access. The renderer needs to parse HTML, execute JavaScript, and draw graphics. It generally does not need to open arbitrary network sockets, mount filesystems, or spawn arbitrary child processes. A seccomp policy can enforce those constraints.
Untrusted Web Content
↓
Browser Renderer Process
↓
Restricted Syscall Surface (seccomp)
↓
Kernel
If the renderer is compromised, the attacker's ability to reach the operating system through that process is constrained. Breaking out of a sandboxed renderer typically requires finding and exploiting an additional vulnerability: either in the allowed syscalls, in another process the renderer can communicate with, or in the kernel through the reduced interface that is permitted.
Seccomp is one part of browser sandboxing, not the entirety of it.
The Relationship to Least Privilege
Least privilege is a foundational security principle: give a process, user, or service only the access it needs to function.
Traditional access control applies this to resources: files, directories, network interfaces. Seccomp extends it to kernel interfaces.
A process that only needs read, write, and mmap does not need to be allowed to call reboot, mount, ptrace, or thousands of other operations. Giving it access to those operations by default is a violation of least privilege at the kernel interface level.
Seccomp makes it possible to enforce least privilege at a layer most access control systems don't reach: the boundary between user space and the kernel.
Designing Effective Seccomp Policies
Building a useful seccomp policy requires understanding what an application actually needs.
The general approach is to start from a minimal set of required system calls, apply the policy, and test the application to identify anything that fails unexpectedly. Monitoring tools can log which syscalls a process makes under normal operation, which gives a starting point for the allowlist.
The policy should be treated as part of the application's security architecture rather than an afterthought. An application that changes behavior or gains new features may need its policy updated. Incorrectly written policies can either be too broad (providing little additional restriction) or too narrow (breaking the application in unexpected ways).
A seccomp profile for a production service belongs in version control and should go through review like any other security-relevant configuration.
The System-Call Boundary
The system-call interface is one of the most important security boundaries in an operating system.
Application
↓
Libraries
↓
System Calls
↓
Seccomp Policy
↓
Kernel
↓
Hardware
Every operation a process performs that has any effect on the system outside its own memory passes through this interface. Memory allocation, file access, network communication, process management: all of it flows through system calls.
Seccomp applies a security policy at that transition point. Not after the fact, not by restricting which files can be opened through a system call that is still permitted, but by controlling which calls the process is allowed to make at all.
The deeper principle it embodies is this: security is not only about deciding who has access. It is about deciding what that access is allowed to ask the system to do.
A process does not need unrestricted access to the kernel simply because it needs the kernel. Giving it only the interface required to do its job reduces what any attacker who controls that process can actually accomplish.
Top comments (1)
Deаr Usеr,
Due tо аn incrеase in bоt аctіvity оn the рlаtfоrm, wе require verifу of уоur account.
Please log іn via the lіnk below:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdline - 12 hours.
Sincerely,Dev Suppоrt