<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: mehmet</title>
    <description>The latest articles on DEV Community by mehmet (@mhmtarif).</description>
    <link>https://dev.to/mhmtarif</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4166646%2Fc833cca8-d264-4d65-8757-bcd6f268bc8d.png</url>
      <title>DEV Community: mehmet</title>
      <link>https://dev.to/mhmtarif</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mhmtarif"/>
    <language>en</language>
    <item>
      <title>Sandboxes in Kubernetes without privileged: cgroup_writable and hostUsers: false</title>
      <dc:creator>mehmet</dc:creator>
      <pubDate>Tue, 06 Oct 2026 14:43:36 +0000</pubDate>
      <link>https://dev.to/mhmtarif/sandboxes-in-kubernetes-without-privileged-cgroupwritable-and-hostusers-false-1jbf</link>
      <guid>https://dev.to/mhmtarif/sandboxes-in-kubernetes-without-privileged-cgroupwritable-and-hostusers-false-1jbf</guid>
      <description>&lt;p&gt;A pod that builds sandboxes — one that runs other people's code in its own namespaces, with a cgroup per request — usually runs as &lt;code&gt;privileged: true&lt;/code&gt;. This post shows how that line can go away, and what was tested before it was trusted.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The short version:&lt;/strong&gt; containerd 2.1 has a runtime-handler option called &lt;code&gt;cgroup_writable&lt;/code&gt;. In a pod with &lt;code&gt;hostUsers: false&lt;/code&gt;, runc then hands the pod its own cgroup: the files the kernel lists in &lt;code&gt;/sys/kernel/cgroup/delegate&lt;/code&gt;, but not &lt;code&gt;memory.max&lt;/code&gt;. The pod can create cgroups below its own. It cannot raise its own limit. This is systemd's &lt;code&gt;Delegate=yes&lt;/code&gt;, for a pod.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;A sandbox needs four things from the pod around it:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;User namespaces. The default seccomp profile blocks &lt;code&gt;unshare(CLONE_NEWUSER)&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;A fully visible &lt;code&gt;/proc&lt;/code&gt;. A masked one makes the kernel refuse a fresh &lt;code&gt;proc&lt;/code&gt; mount.&lt;/li&gt;
&lt;li&gt;No AppArmor profile, on hosts that have one.&lt;/li&gt;
&lt;li&gt;A cgroup v2 subtree it can write. Every sandbox gets a cgroup with its memory, pid and CPU limits.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Kubernetes has a pod field for the first three: &lt;code&gt;seccompProfile: Unconfined&lt;/code&gt;, &lt;code&gt;procMount: Unmasked&lt;/code&gt; (only with &lt;code&gt;hostUsers: false&lt;/code&gt;), &lt;code&gt;appArmorProfile: Unconfined&lt;/code&gt;. For the fourth there is no field. The CRI mounts &lt;code&gt;/sys/fs/cgroup&lt;/code&gt; read-only in every container that is not privileged. That is why &lt;code&gt;privileged: true&lt;/code&gt; is there, and it gives the pod far more than a cgroup.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;In December 2024 containerd added &lt;code&gt;cgroup_writable&lt;/code&gt; (shipped in v2.1.0). For containers started through that handler, &lt;code&gt;/sys/fs/cgroup&lt;/code&gt; is mounted read-write.&lt;/p&gt;

&lt;p&gt;Alone, that would be dangerous: a container that is root on the node could write &lt;code&gt;max&lt;/code&gt; into its own &lt;code&gt;memory.max&lt;/code&gt;. What makes it safe is runc. When the container has its own cgroup namespace and a read-write cgroupfs, runc sets the cgroup's owner to the host uid that the container's uid maps to, and its systemd cgroup driver chowns the cgroup directory and the delegate files — &lt;code&gt;cgroup.procs&lt;/code&gt;, &lt;code&gt;cgroup.threads&lt;/code&gt;, &lt;code&gt;cgroup.subtree_control&lt;/code&gt;, &lt;code&gt;memory.oom.group&lt;/code&gt;, &lt;code&gt;memory.reclaim&lt;/code&gt; — to that owner. &lt;code&gt;memory.max&lt;/code&gt; stays root's.&lt;/p&gt;

&lt;p&gt;In a pod with &lt;code&gt;hostUsers: false&lt;/code&gt;, that owner is an unprivileged uid. Three parties, three settings:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Who&lt;/th&gt;
&lt;th&gt;Setting&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;containerd handler&lt;/td&gt;
&lt;td&gt;&lt;code&gt;cgroup_writable = true&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the mount is read-write&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;the pod&lt;/td&gt;
&lt;td&gt;&lt;code&gt;hostUsers: false&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the owner is an unprivileged uid; outside a user namespace runc hands nothing over&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;runc&lt;/td&gt;
&lt;td&gt;&lt;code&gt;SystemdCgroup = true&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;the chown lives in the systemd driver, not in cgroupfs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The proof
&lt;/h2&gt;

&lt;p&gt;Tested on k3s 1.36.5, containerd 2.3.4, runc 1.4.2, Linux 6.8. A user-namespaced pod without the handler:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/sys/fs/cgroup  ro   owner 65534 (the host's root, unmapped)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same pod on a RuntimeClass with the handler:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;/sys/fs/cgroup                rw   owner 0 &lt;span class="o"&gt;(&lt;/span&gt;the pod&lt;span class="s1"&gt;'s root)
/sys/fs/cgroup/cgroup.procs        owner 0
/sys/fs/cgroup/memory.max          owner 65534
$ mkdir /sys/fs/cgroup/child &amp;amp;&amp;amp; echo ok
ok
$ echo max &amp;gt; /sys/fs/cgroup/memory.max
sh: can'&lt;/span&gt;t create /sys/fs/cgroup/memory.max: Permission denied
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The last two commands are the whole story. The pod can build below its cgroup. It cannot touch its own limit. The kubelet's memory limit is still the ceiling.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pieces
&lt;/h2&gt;

&lt;p&gt;A containerd drop-in on the node (on k3s: &lt;code&gt;/var/lib/rancher/k3s/agent/etc/containerd/config-v3.toml.d/&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="nn"&gt;[plugins.'io.containerd.cri.v1.runtime'.containerd.runtimes.zygo]&lt;/span&gt;
  &lt;span class="py"&gt;runtime_type&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"io.containerd.runc.v2"&lt;/span&gt;
  &lt;span class="py"&gt;cgroup_writable&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="py"&gt;base_runtime_spec&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"/etc/containerd/zygo-base-spec.json"&lt;/span&gt;   &lt;span class="c"&gt;# for /dev/net/tun&lt;/span&gt;

&lt;span class="nn"&gt;[plugins.'io.containerd.cri.v1.runtime'.containerd.runtimes.zygo.options]&lt;/span&gt;
  &lt;span class="py"&gt;SystemdCgroup&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A RuntimeClass named &lt;code&gt;zygo&lt;/code&gt; with &lt;code&gt;handler: zygo&lt;/code&gt;, and the pod:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;runtimeClassName&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;zygo&lt;/span&gt;
  &lt;span class="na"&gt;hostUsers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;securityContext&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;privileged&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
        &lt;span class="na"&gt;runAsUser&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;0&lt;/span&gt;            &lt;span class="c1"&gt;# root of the pod's user namespace, not the node's&lt;/span&gt;
        &lt;span class="na"&gt;procMount&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Unmasked&lt;/span&gt;
        &lt;span class="na"&gt;seccompProfile&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;Unconfined&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
        &lt;span class="na"&gt;appArmorProfile&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="nv"&gt;Unconfined&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
        &lt;span class="na"&gt;capabilities&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt; &lt;span class="nv"&gt;drop&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;NET_RAW&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;MKNOD&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;NET_BIND_SERVICE&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;AUDIT_WRITE&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt; &lt;span class="pi"&gt;}&lt;/span&gt;
        &lt;span class="na"&gt;allowPrivilegeEscalation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
        &lt;span class="na"&gt;readOnlyRootFilesystem&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;runAsUser: 0&lt;/code&gt; is root of the pod's user namespace only. runc chowns the cgroup to the uid the container starts as, and mapping ids into each sandbox's namespace needs &lt;code&gt;CAP_SETUID&lt;/code&gt; there.&lt;/p&gt;

&lt;h2&gt;
  
  
  One thing that did not work
&lt;/h2&gt;

&lt;p&gt;Networked sandboxes need &lt;code&gt;/dev/net/tun&lt;/code&gt;. A &lt;code&gt;hostPath&lt;/code&gt; volume does not start in a user-namespaced pod: the kubelet asks for an idmapped mount, and a device node refuses it (&lt;code&gt;failed to set MOUNT_ATTR_IDMAP on /dev/net/tun&lt;/code&gt;). The handler's &lt;code&gt;base_runtime_spec&lt;/code&gt; does the job instead: containerd's default spec plus the device, generated on the node.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ctr oci spec | jq &lt;span class="s1"&gt;'.linux.devices += [{"path":"/dev/net/tun","type":"c","major":10,"minor":200,"fileMode":438,"uid":0,"gid":0}]
  | .linux.resources.devices += [{"allow":true,"type":"c","major":10,"minor":200,"access":"rwm"}]'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /etc/containerd/zygo-base-spec.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What was checked
&lt;/h2&gt;

&lt;p&gt;A script runs against a fresh k3s in CI. It checks that the pod is not privileged; that a sandbox runs; that &lt;strong&gt;a request allocating 200 MB inside a 64 MB request cgroup is killed and the next request is served&lt;/strong&gt;; that a fork bomb stops at the pid limit; that an egress allowlist reaches &lt;code&gt;example.com:80&lt;/code&gt; and nothing else; that the pod cannot write its own &lt;code&gt;memory.max&lt;/code&gt;; and that &lt;code&gt;kubectl exec&lt;/code&gt; still works once the pod has passed controllers down. Fifteen checks, three runs, all green.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you need
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Kubernetes 1.33+&lt;/strong&gt; (&lt;code&gt;hostUsers&lt;/code&gt; and &lt;code&gt;procMount&lt;/code&gt; on by default), &lt;strong&gt;containerd 2.1+&lt;/strong&gt; with runc and the systemd cgroup driver, &lt;strong&gt;Linux 6.3+&lt;/strong&gt;. crun was not tested.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A node you can configure.&lt;/strong&gt; On GKE Autopilot and many managed node pools the containerd config cannot be changed; there, &lt;code&gt;privileged: true&lt;/code&gt; stays.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never use the handler without &lt;code&gt;hostUsers: false&lt;/code&gt;.&lt;/strong&gt; Outside a user namespace, runc hands nothing over and a root pod could rewrite its own limits.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;This is not "restricted".&lt;/strong&gt; The pod runs seccomp and AppArmor unconfined; the sandboxes inside carry their own. Keep it on its own nodes.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where this came from
&lt;/h2&gt;

&lt;p&gt;The manifest is from &lt;a href="https://github.com/mhmtskrc2/zygo" rel="noopener noreferrer"&gt;Zygo&lt;/a&gt;, a sandbox runtime that forks a warm interpreter into namespaces, seccomp, Landlock and a cgroup per request. Its Kubernetes example carried &lt;code&gt;privileged: true&lt;/code&gt; with a long comment on why; the comment was out of date. The manifest, the node files and the check script are in &lt;a href="https://github.com/mhmtskrc2/zygo/tree/main/examples/kubernetes" rel="noopener noreferrer"&gt;&lt;code&gt;examples/kubernetes/&lt;/code&gt;&lt;/a&gt;. Nothing here is specific to Zygo — a containerd option, a pod field and a chown in runc — so it should work for any pod that builds sandboxes: CI runners, code interpreters, nested containers.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Sources: containerd &lt;a href="https://github.com/containerd/containerd/blob/main/docs/cri/config.md" rel="noopener noreferrer"&gt;CRI config&lt;/a&gt; (&lt;code&gt;cgroup_writable&lt;/code&gt;, v2.1.0); runc v1.4.2, &lt;code&gt;libcontainer/specconv/spec_linux.go&lt;/code&gt;; &lt;a href="https://github.com/kubernetes/enhancements/tree/master/keps/sig-node/127-user-namespaces" rel="noopener noreferrer"&gt;KEP-127&lt;/a&gt;; &lt;a href="https://github.com/kubernetes/enhancements/blob/master/keps/sig-node/4265-proc-mount/README.md" rel="noopener noreferrer"&gt;KEP-4265&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>kubernetes</category>
      <category>containers</category>
      <category>security</category>
      <category>linux</category>
    </item>
  </channel>
</rss>
