Something is off with a Pod. The app logs show nothing that looks like an error.
At this point I want to get inside and look around: open a shell with kubectl exec, check processes with ps, send a request to the app with curl localhost. The usual routine.
But the shell won't start.
$ kubectl exec app -c app -- sh
error: Internal error occurred: Internal error occurred: error executing command in container: failed to exec in container: failed to start exec "5559c54c81498e2d06cfaf7704ee420b7a48a48ef1531e34307aa515f53083b3": OCI runtime exec failed: exec failed: unable to start container process: exec: "sh": executable file not found in $PATH
The last part, exec: "sh": executable file not found in $PATH, is the answer: this container has no sh. It has no ls or cat either.
The container runs from a distroless image. A distroless image contains only the app's binary and the minimum files it needs to run. It leaves out the shell and the package manager so that there are no tools for an attacker to use. As a result, even during an incident, you can't run commands inside the container to investigate.
Kubernetes has a mechanism for this situation. You can add a temporary container to a running Pod. This is called an ephemeral container, and you add one with the kubectl debug command.
You leave the app container untouched and add a separate container to the same Pod, one that has curl, ss, and tcpdump. Containers in a Pod share the network, so running curl localhost:3000 from the added container reaches the app directly.
The official docs cover the basic usage of kubectl debug, so this article covers what you can see once you're inside.
Once you're inside, you quickly run into cases where you can't get the information you want.
- You try to read the app's config file, and nothing exists at that path
- You check the CPU limit in
/sys/fs/cgroup/cpu.max, and instead of the app's50000 100000you get your own unlimitedmax 100000 - There's a way around to see the app's files, but without
SYS_PTRACEyou getPermission denied
Even within the same Pod, how far you can reach into the app is determined by Linux namespaces, and whether you can read what you reach is determined by permissions. This article doesn't walk through tracking down the cause of an incident. It checks, one item at a time against real output, what you can and can't see once inside along these two axes, and shows a workaround for each thing you can't see.
Test setup
I ran everything on a kind v0.33.0 cluster on a GitHub Actions x86_64 runner (ubuntu-24.04, kernel 6.17). Kubernetes is v1.37.0 and the container runtime is containerd v2.3.4. The target app is a small HTTP server written in Go, built on gcr.io/distroless/static-debian12:nonroot. For the debug image I used nicolaka/netshoot:v0.16. All output in this article is pasted as-is from this environment.
kind is a tool that runs a Kubernetes cluster using Docker containers as nodes. I used a cluster with a single node (ded-control-plane).
I put one Pod, app, on this cluster and added containers to it with kubectl debug to investigate.
Node ded-control-plane (kind)
└── Pod app
├── app distroless HTTP server (:3000, uid 65532)
├── sidecar HTTP server from the same image (:4000)
├── dbg-notarget added with kubectl debug (no --target)
├── dbg-default added with kubectl debug (--target=app)
└── dbg-sysadmin added with kubectl debug (--target=app --profile=sysadmin)
The app container is the one under investigation. The distroless :nonroot image sets the runtime user to uid 65532 (nonroot), and this Pod doesn't override it with runAsUser, so app runs as uid 65532. Two of its settings matter in later sections.
- The CPU limit is
500m(half a core) - The ConfigMap's
app.jsonis mounted at/config
sidecar is a regular container defined in the Pod from the start, running the same image as app on a different port. It's there to compare what containers in a Pod share.
The three containers whose names start with dbg- are ephemeral containers added later with kubectl debug. They differ only in the options passed.
--target=app makes the added container see the same set of processes as the app container. With it, app's processes show up in ps.
--profile selects the permissions the added container gets. If you don't specify it, you get general, which adds one capability for inspecting other processes (SYS_PTRACE). dbg-notarget and dbg-default use this profile. Specifying sysadmin gives you a privileged container, meaning one with nearly every permission. A later section shows measured differences between the profiles.
The reproduction code is here.
https://github.com/shinagawa-web/show-your-work/tree/main/distroless-ephemeral-debug
Background: container isolation is set per namespace
Linux can partition what a process sees (the network, process IDs, the filesystem, and so on), each one separately. Each of these partitions is called a namespace. A container is built from a combination of these namespaces.
The key point is that each namespace is independent. You can't describe a container as simply "isolated". It's normal for a container to share the network with its neighbor while having a separate filesystem. This explains why, from a container added with kubectl debug, curl localhost reaches the app while the app's config file is not visible. Both happen at the same time.
You can check how they're actually split from the node. Each process has one link per namespace under /proc/<pid>/ns/. Reading /proc/<pid>/ns/net with readlink, for example, returns a string like net:[4026532670]. The number in brackets identifies the namespace, and processes with the same number are in the same namespace.
To do this you need each container's PID. On the node ded-control-plane, I used crictl (a command that talks to the container runtime directly) and took each container's process PID on the node from .info.pid in crictl inspect <container ID>. For each PID I ran readlink on every link under /proc/<pid>/ns/ and put the results on one line per container. That gives the table below, trimmed to the rows and columns this article needs. I dropped the ipc, uts, user, and time columns because this article doesn't use them.
container node_pid net pid mnt cgroup
node-init 1 net:[4026532365] pid:[4026532362] mnt:[4026532359] cgroup:[4026532364]
app 4082 net:[4026532670] pid:[4026532736] mnt:[4026532735] cgroup:[4026532737]
sidecar 4118 net:[4026532670] pid:[4026532739] mnt:[4026532738] cgroup:[4026532740]
dbg-notarget 4193 net:[4026532670] pid:[4026532742] mnt:[4026532741] cgroup:[4026532743]
dbg-default 4247 net:[4026532670] pid:[4026532736] mnt:[4026532744] cgroup:[4026532745]
dbg-sysadmin 4863 net:[4026532670] pid:[4026532736] mnt:[4026532754] cgroup:[4026532364]
node_pid is the PID on the node from crictl, and the columns to its right are the readlink results. node-init is PID 1 on the node, included for comparison.
net has the same number for every container. For pid, only dbg-default and dbg-sysadmin have the same number as app. mnt is different for every container. For cgroup, only dbg-sysadmin has the same number as the node.
The next table summarizes these results per namespace and adds what each namespace separates and what a debug container can see.
| namespace | What it separates | Between containers in a Pod | Ephemeral container | Result |
|---|---|---|---|---|
| network | IP addresses, ports, sockets | Shared | Shared |
curl localhost and ss show the app's traffic |
| PID | Process IDs and which processes are visible | Separate (can be shared with shareProcessNamespace) |
Shared with the target when --target is set, separate otherwise |
With --target, the target's processes show up in ps
|
| mount | How the filesystem looks | Separate | Separate | The other container's files aren't directly visible |
| cgroup | How CPU and memory limits look | Separate | Separate (only sysadmin shares it with the node) |
/sys/fs/cgroup shows your own values (except with sysadmin) |
This table covers how far you can reach into the app. Whether you can read what you reach depends on permissions, which I cover in "Permissions: SYS_PTRACE and profiles".
Adding an ephemeral container
For the debug image I use nicolaka/netshoot. It bundles network troubleshooting tools such as curl, ss, tcpdump, and dig, along with basic commands like ps.
To get an interactive shell, run this.
$ kubectl debug -it app --image=nicolaka/netshoot:v0.16 --target=app -c dbg-it -- bash
Targeting container "app". If you don't see processes from this container it may be because the container runtime doesn't support this feature.
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
Two warnings appear before the shell opens.
The first says that --target only works if the container runtime supports it. In this article I confirmed it works on containerd v2.3.4. I don't cover what things look like on a runtime without support.
The second says that every command you type in this session, and its output, is recorded in the container logs. For example, if you print /proc/1/environ as in "Environment and arguments: read from /proc/1" below, the app's environment variables end up in these logs, secrets included.
--target changes what ps shows
In the table in the background section, the PID row was the only one that changed with --target. Let's check this with ps.
For this test I wanted to run commands in the same container repeatedly and compare the output. So instead of entering with -it, I kept each container running with sleep infinity and ran commands with kubectl exec -c <container name>. I added dbg-notarget without --target and dbg-default with --target=app as follows. --image-pull-policy=Never is there because the image was preloaded onto the kind node.
$ kubectl debug app --image=nicolaka/netshoot:v0.16 --image-pull-policy=Never -c dbg-notarget -- sleep infinity
$ kubectl debug app --image=nicolaka/netshoot:v0.16 --image-pull-policy=Never -c dbg-default --target=app -- sleep infinity
Targeting container "app". If you don't see processes from this container it may be because the container runtime doesn't support this feature.
I ran ps auxf in each.
$ kubectl exec app -c dbg-notarget -- ps auxf
PID USER TIME COMMAND
1 root 0:00 sleep infinity
14 root 0:00 ps auxf
$ kubectl exec app -c dbg-default -- ps auxf
PID USER TIME COMMAND
1 65532 0:00 /server -http=:3000
29 root 0:00 sleep infinity
42 root 0:00 ps auxf
Without --target, all you see is your own sleep infinity and the ps you just ran. The app's process is nowhere to be found.
With --target=app, the app's /server shows up as PID 1, and your own sleep infinity becomes PID 29. The 65532 in the USER column is the uid of the distroless nonroot user.
The app starts as PID 1 inside its container, so it also appears as PID 1 from a container that joins its PID namespace with --target. That's why later sections write /proc/1/.... However, if the app is launched through an init process like tini, or if the Pod has shareProcessNamespace enabled, the app may not be PID 1. Check with ps first.
dbg-notarget, which has no --target, also has the capability for inspecting other processes (SYS_PTRACE). But it can't see the app's process at all, so there's nothing to use it on. Every step below that uses /proc/1/root or /proc/1/environ assumes you entered with --target.
What you can see directly
From here on, I check the background table against real output. First come the things an ephemeral container reaches directly. The network namespace is shared across the whole Pod, and the PID namespace is shared with app through --target. Unless stated otherwise, the results are from inside dbg-default, which was added with --target=app.
dbg-default has SYS_PTRACE, so everything read in this section was readable except /proc/1/stack. Some of these become unreadable without SYS_PTRACE. I cover that in "Permissions: SYS_PTRACE and profiles".
Network: use it directly
The network row in the table said "Shared". From dbg-default, you see the same network as the app.
$ ss -tanp
State Recv-Q Send-Q Local Address:Port Peer Address:PortProcess
LISTEN 0 4096 *:3000 *:* users:(("server",pid=1,fd=4))
LISTEN 0 4096 *:4000 *:*
:3000 is the port the app (server, pid 1) listens on. :4000 is the sidecar's port. You can see that something is listening on it, but not which process.
This also matches the table. The network is shared across the whole Pod, so the sidecar's socket is visible. The PID namespace, on the other hand, is shared only with app, so the sidecar's process isn't visible. ss finds a socket's owner by looking through the processes under /proc, so it can't show the name of a process that isn't visible.
localhost reaches the app as well.
$ curl -sv -m 3 localhost:3000/health
...
< HTTP/1.1 200 OK
< Date: Sun, 27 Sep 2026 02:28:19 GMT
< Content-Length: 3
< Content-Type: text/plain; charset=utf-8
<
{ [3 bytes data]
ok
* Connection #0 to host localhost:3000 left intact
The response is 200 OK with the body ok, so the app answers requests from localhost. The ... part is the connection setup, which I left out.
You can capture packets too. I ran tcpdump while sending curl requests in the background, and stopped after the first 2 packets, which show the SYN and SYN-ACK of one connection.
$ (for i in 1 2 3; do sleep 1; curl -s -m 2 localhost:3000/health >/dev/null; done) & tcpdump -i any -nn -c 2 port 3000; rc=$?; wait; exit $rc
tcpdump: WARNING: any: That device doesn't support promiscuous mode
(Promiscuous mode not supported on the "any" device)
tcpdump: verbose output suppressed, use -v[v]... for full protocol decode
listening on any, link-type LINUX_SLL2 (Linux cooked v2), snapshot length 262144 bytes
02:28:20.415650 lo In IP6 ::1.43214 > ::1.3000: Flags [S], seq 838274280, win 65476, options [mss 65476,sackOK,TS val 235319763 ecr 0,nop,wscale 10], length 0
02:28:20.415666 lo In IP6 ::1.3000 > ::1.43214: Flags [S.], seq 4234760723, ack 838274281, win 65464, options [mss 65476,sackOK,TS val 235319763 ecr 235319763,nop,wscale 10], length 0
2 packets captured
20 packets received by filter
0 packets dropped by kernel
Both captured packets are IPv6 traffic over lo (loopback). The first is a connection request (Flags [S], SYN) from port 43214 on ::1 to the app's :3000. The second is the app's reply (Flags [S.], SYN-ACK). Traffic to the app can be captured from inside the debug container.
Name resolution works too. I looked up the name of the Service app against the nameserver listed in dbg-default's own /etc/resolv.conf.
$ dig @$(awk '/^nameserver/{print $2; exit}' /etc/resolv.conf) app.default.svc.cluster.local +short
10.96.248.99
The returned 10.96.248.99 is the ClusterIP of the Service app. The app's environment variable APP_SERVICE_HOST holds the same value.
Process state: read from /proc/1
As shown in "--target changes what ps shows", the app appears as PID 1 when you enter with --target=app. Its process state is readable under /proc/1/.
$ grep -E '^(Name|State|Uid|Threads|VmRSS):' /proc/1/status
Name: server
State: S (sleeping)
Uid: 65532 65532 65532 65532
VmRSS: 7256 kB
Threads: 4
State is S (sleeping), meaning the process is blocked waiting for something. Threads shows 4 threads, and VmRSS, the physical memory actually in use, is 7256 kB. Uid is 65532 in every field because the process runs as the nonroot user.
The list of open files is in /proc/1/fd.
$ ls -l /proc/1/fd
total 0
lrwx------ 1 65532 65532 64 Sep 27 02:28 0 -> /dev/null
l-wx------ 1 65532 65532 64 Sep 27 02:28 1 -> pipe:[45581]
l-wx------ 1 65532 65532 64 Sep 27 02:28 2 -> pipe:[45582]
lr-x------ 1 65532 65532 64 Sep 27 02:28 3 -> /sys/fs/cgroup/cpu.max
lrwx------ 1 65532 65532 64 Sep 27 02:28 4 -> socket:[45610]
lrwx------ 1 65532 65532 64 Sep 27 02:28 5 -> anon_inode:[eventpoll]
lrwx------ 1 65532 65532 64 Sep 27 02:28 6 -> anon_inode:[eventfd]
fd 4 is a socket, and it corresponds to fd=4 (:3000) in the earlier ss -tanp output.
For slow operations such as receiving from the network or reading from disk, a program makes a system call and blocks inside the kernel (the core of the OS) until the call completes. /proc/1/wchan (short for wait channel) holds the name of the kernel function where the process is currently blocked, so it tells you what the process is waiting for.
$ cat /proc/1/wchan; echo
ep_poll
In this run it was ep_poll. When I entered the same Pod with --profile=sysadmin, it returned futex_do_wait. The value depends on when you read it, so read it a few times instead of judging from a single read.
-
ep_poll: waiting for data to arrive on one of its connections. This is the idle state of waiting for requests -
futex_do_wait: blocked while synchronizing with another thread of the same program. In Go apps, idle threads are often in this state
For more detail on what the process is waiting for, /proc/1/stack prints the call stack inside the kernel. However, it's not readable from dbg-default.
$ cat /proc/1/stack
cat: read error: Permission denied
Among the profiles I tried, only the privileged dbg-sysadmin could read it.
$ cat /proc/1/stack
[<0>] futex_do_wait+0x3a/0x70
[<0>] __futex_wait+0x99/0x100
[<0>] futex_wait+0x72/0x120
[<0>] do_futex+0xc9/0x190
[<0>] __x64_sys_futex+0x125/0x1f0
[<0>] x64_sys_call+0x1d1d/0x20d0
[<0>] do_syscall_64+0x7b/0xb40
[<0>] entry_SYSCALL_64_after_hwframe+0x76/0x7e
This is the call path inside the kernel, read from bottom to top. The bottom is the entry point where the app entered the kernel, and the top is where it's blocked right now. In this output, the app is waiting in futex (a system call for synchronizing threads) and is blocked in futex_do_wait at the top. Numbers like +0x3a/0x70 are offsets within the function, and you can usually ignore them.
The top entry, futex_do_wait, is the same as the wchan value read from the same dbg-sysadmin. stack also shows the path that led there. Still, if all you need is what the process is waiting for, wchan is usually enough.
Environment and arguments: read from /proc/1
The app's environment variables at startup are in /proc/1/environ. The entries are separated by NUL characters, so replace them with newlines to read them.
$ tr '\0' '\n' < /proc/1/environ
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
HOSTNAME=app
SSL_CERT_FILE=/etc/ssl/certs/ca-certificates.crt
APP_MODE=production
APP_UPSTREAM={"upstream": "http://backend.default.svc:8080", "timeout_ms": 250}
...
After the ..., the output continues with Service addresses and other variables that Kubernetes injects automatically.
The startup arguments are in /proc/1/cmdline.
$ tr '\0' ' ' < /proc/1/cmdline; echo
/server -http=:3000
The only argument is -http=:3000, the same command line that ps showed.
What you can't see directly
Next come the things the table marks "Separate", which you can't reach just by entering: mount and cgroup. Both have a workaround through /proc/1/root.
Files: read from /proc/1/root
In the introduction I wrote that when you try to read the app's config file, nothing exists at that path. Here is that case. The app mounts a ConfigMap at /config, but reading the same path from dbg-default finds nothing.
$ cat /config/app.json
cat: can't open '/config/app.json': No such file or directory
The mount row in the table said "Separate". The / in dbg-default is the netshoot image's filesystem, which is separate from the app's.
This is where /proc/1/root comes in. /proc/<pid>/root is a link that points to "/ as seen by that process", and following it takes you straight into the filesystem the app sees. You can read the contents of an image that has no shell and no ls with the tools on your side.
$ ls /proc/1/root/
bin
boot
config
dev
etc
home
lib
proc
...
root
run
sbin
scratch
server
sys
tmp
usr
var
config is where the ConfigMap is mounted, scratch is where the Pod spec mounts an emptyDir, and server is the app's binary. The ConfigMap's contents are readable too.
$ cat /proc/1/root/config/app.json
{"upstream": "http://backend.default.svc:8080", "timeout_ms": 250}
The config the app reads is now readable from the debug container.
If you suspect name resolution, you can read the resolv.conf the app actually uses in the same way.
$ cat /proc/1/root/etc/resolv.conf
search default.svc.cluster.local svc.cluster.local cluster.local urp3p5u2pksefpdlab2z34ww2g.gx.internal.cloudapp.net
nameserver 10.96.0.10
options ndots:5
10.96.0.10 in nameserver is where the app sends its queries. search lists the domains that are appended to short names in order. ndots:5 means that for a name with fewer than 5 dots, the search domains are appended and tried first. When a name doesn't resolve, this tells you which names the app actually queries. The last search entry, …cloudapp.net, isn't a cluster domain. It comes from the host side's configuration.
cgroup: read from /proc/1/root
A cgroup is the mechanism that applies CPU and memory limits per container. The limit values are readable from files under /sys/fs/cgroup. The app's CPU limit is 500m (half a core), so cpu.max should contain 50000 100000 (up to 50,000 microseconds out of every 100,000 microseconds).
From dbg-default, I checked where my own cgroup is, that /sys/fs/cgroup lists cgroup files, and the value of cpu.max, in one go. The result looks like this.
$ cat /proc/self/cgroup; ls /sys/fs/cgroup | head -n 8; cat /sys/fs/cgroup/cpu.max
0::/
cgroup.controllers
cgroup.events
cgroup.freeze
cgroup.kill
cgroup.max.depth
cgroup.max.descendants
cgroup.pressure
cgroup.procs
max 100000
max 100000 means "no limit", and it's not the app's value. This is the case from the introduction where you get your own value. What you see is the debug container's own cgroup, which has no limit.
The cgroup row in the table said "Separate". dbg-default has a cgroup namespace of its own, and its own cgroup appears at the top of /sys/fs/cgroup. The first line, 0::/, means "my cgroup is at the top of the hierarchy as I see it".
To read the app's values, go through /proc/1/root as in "Files: read from /proc/1/root". Reading /sys/fs/cgroup as the app sees it gives the app's own values.
$ cat /proc/1/root/sys/fs/cgroup/cpu.max
50000 100000
The output is 50000 100000, matching the app's 500m limit. This path goes through /proc/1/root, so whether you can read it depends on having SYS_PTRACE. It was readable with general (no --profile) and sysadmin, and returned Permission denied with legacy, baseline, and netadmin. I explain why in "Permissions: SYS_PTRACE and profiles".
The exception is sysadmin. dbg-sysadmin shares the cgroup namespace with the node, so the app's cgroup path shown in /proc/1/cgroup exists at the same path under its own /sys/fs/cgroup.
$ ls -d /sys/fs/cgroup$(cut -d: -f3 /proc/1/cgroup)
/sys/fs/cgroup/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-pod719135d1_ee24_4f46_948a_440e1c6f4f77.slice/cri-containerd-698b44cb65ab91fba808118cba48f23ab4761ddd8a3e05977a04c4323bcef648.scope
This reaches the app's cgroup directory directly, without going through /proc/1/root.
If you have permission to get onto the node, you can also read the app's cgroup directly from the node side. I cover that later in "Reading cgroups from the node".
Permissions: SYS_PTRACE and profiles
The reads of /proc/1/root and /proc/1/environ so far worked because dbg-default has SYS_PTRACE. The table in the background section only covers how far you can reach into the app. Whether you can read what you reach depends on the permissions in this section.
SYS_PTRACE: without it, you can't read the contents
SYS_PTRACE is what decides between reading the contents of /proc/1 and getting Permission denied.
When you read /proc/1/root, /proc/1/environ, /proc/1/maps, the link targets in /proc/1/fd, and similar files, the kernel checks whether the reading process is allowed to inspect the contents of the other process. dbg-default runs as root (uid 0 / gid 0). The app runs as uid 65532, and its gid is also 65532 (you can see this in the owner columns, 65532 65532, of the earlier ls -l /proc/1/fd). When your uid and gid differ from the target's like this, inspecting it requires SYS_PTRACE.
This isn't about file permissions. You're root, and you even have the capability to bypass file permissions (DAC_OVERRIDE), but that doesn't get you past this check.
To confirm this, I also entered with --profile=legacy, a profile without SYS_PTRACE, ran the same commands as with default, and compared the results. This container is dbg-legacy, added to the same Pod app with only --profile changed. legacy is deprecated, and kubectl prints a warning when you add it saying that it's scheduled for removal in v1.39. baseline, which also lacks SYS_PTRACE, gave the same results as legacy.
I compared 9 items that read the contents of the app's process. They are the process column of ss -tanp, the link targets in /proc/1/fd, /proc/1/maps, /proc/1/environ, the listing of /proc/1/root/, and reads of app.json, resolv.conf, cpu.max, and cpu.stat through /proc/1/root/.
All 9 items failed with legacy, and all 9 passed with default.
Here are two examples of the failures after entering app with legacy.
$ cat /proc/1/root/config/app.json
cat: can't open '/proc/1/root/config/app.json': Permission denied
$ ls -l /proc/1/fd
total 0
ls: /proc/1/fd/0: cannot read link: Permission denied
lrwx------ 1 65532 65532 64 Sep 27 02:28 0
...
Both are Permission denied. ls -l /proc/1/fd prints the list but can't read the link targets.
The users:(...) column disappears from ss -tanp.
$ ss -tanp
State Recv-Q Send-Q Local Address:Port Peer Address:PortProcess
LISTEN 0 4096 *:3000 *:*
LISTEN 0 4096 *:4000 *:*
...
The :3000 line has no users:(("server",pid=1,fd=4)) either, so you can't tell which process owns the socket.
wchan doesn't produce an error, which makes this case easy to miss. Reading it after entering app with legacy returned 0 instead of futex_do_wait or ep_poll.
$ cat /proc/1/wchan; echo
0
If wchan is 0, first check whether you have SYS_PTRACE.
On the other hand, /proc/1/status, /proc/1/cmdline, and the listing of /proc/1/fd itself were readable even with legacy.
SYS_PTRACE is the capability for inspecting another process, and it's useless if that process isn't visible. --target makes the target visible, and SYS_PTRACE grants permission to inspect it. The commands so far work only with both in place.
uid and gid: when both match, reads worked without SYS_PTRACE
There's another way past this kernel check. In this test, when both the uid and the gid of my process matched the target's, reads worked without SYS_PTRACE. However, every target I tried drops all capabilities with capabilities.drop: ["ALL"]. The kernel check also has a condition on the target's capabilities, and I didn't test targets that keep any capabilities.
To confirm this, I prepared two Pods in addition to app, with only runAsUser and runAsGroup changed. app-uid0 runs the app as uid 0 / gid 65532, and app-uid0-gid0 runs it as uid 0 / gid 0.
Reading each app's /proc/<pid>/status from the node gave the following. The first block is app-uid0, and the second is app-uid0-gid0.
$ grep -E '^(Uid|Gid|Groups|CapInh|CapPrm|CapEff|CapBnd|CapAmb|NoNewPrivs|Seccomp):' /proc/7049/status
Uid: 0 0 0 0
Gid: 65532 65532 65532 65532
...
CapEff: 0000000000000000
...
$ grep -E '^(Uid|Gid|Groups|CapInh|CapPrm|CapEff|CapBnd|CapAmb|NoNewPrivs|Seccomp):' /proc/7042/status
Uid: 0 0 0 0
Gid: 0 0 0 0
...
CapEff: 0000000000000000
...
app-uid0 has uid 0 and gid 65532, and app-uid0-gid0 has uid 0 and gid 0. Both have CapEff set to all zeros, so both run with every capability dropped.
I entered each one with legacy and ran the same commands. The results for app are included for comparison.
| Target app | uid / gid | Result with legacy |
|---|---|---|
app |
65532 / 65532 | All 9 items failed |
app-uid0 |
0 / 65532 | All 9 items failed |
app-uid0-gid0 |
0 / 0 | All passed |
Matching only the uid (app-uid0) isn't enough. The reads pass only when the gid matches as well (app-uid0-gid0).
wchan follows the same condition. Entering app-uid0-gid0 with legacy returned futex_do_wait rather than 0. So the 0 seen in "SYS_PTRACE: without it, you can't read the contents" can be read as the value the kernel returns when this check fails.
The distroless nonroot image uses 65532 for both uid and gid. As long as you enter as root, both the uid and the gid differ from the app's, so the quickest route is to enter with a profile that includes SYS_PTRACE.
Profiles: only general and sysadmin can read /proc/1/root
The capabilities the debug container gets are determined by --profile, and the profile determines which commands fail.
What differs between profiles is the securityContext, the part of the spec that sets things like a container's capabilities. I also added legacy, general, baseline, and netadmin to the same Pod app as dbg-default and dbg-sysadmin, changing only --profile, and ran the same commands to compare them. restricted requires an image that runs as non-root, and it didn't start with netshoot, which runs as root, so I left it out of the comparison. Here are the results.
| profile | securityContext | Capabilities | /proc/1/root |
iptables |
|---|---|---|---|---|
general (same as no --profile) |
capabilities.add: [SYS_PTRACE] |
Base + SYS_PTRACE
|
Readable | Fails |
| legacy / baseline | None | Base only | Not readable | Fails |
| netadmin | capabilities.add: [NET_ADMIN, NET_RAW] |
Base + NET_ADMIN
|
Not readable | Works |
| sysadmin | privileged: true |
Nearly all | Readable | Works |
"Base" in the table is the set of capabilities the container runtime (containerd) grants when the securityContext specifies nothing.
Without NET_ADMIN, iptables fails with the misleading error Permission denied (you must be root), even when it runs as root. tcpdump worked with every profile in the table, because NET_RAW, which packet capture needs, is in the base set from the start. netadmin adds NET_ADMIN and NET_RAW, but NET_RAW is already there, so the only real addition is NET_ADMIN. That's why the Capabilities column shows "Base + NET_ADMIN".
Reading cgroups from the node
As shown in "cgroup: read from /proc/1/root", an ephemeral container can read the app's cgroup values through /proc/1/root/sys/fs/cgroup/..., and with sysadmin you can follow them directly from /sys/fs/cgroup. Looking from the node is the way out when neither of these is available. Here are two such cases.
- Pod Security settings or similar policies on the app's Kubernetes namespace allow neither
SYS_PTRACEnor privileged containers - You want to look across multiple Pods on the node
Even in the first case, starting the node debug Pod in a different Kubernetes namespace that allows privileged Pods should let you get onto the node. I didn't test that far in this article, though.
Enter the node with kubectl debug node
kubectl debug node/<node name> starts a debug Pod on the specified node and attaches you to it.
$ kubectl debug node/ded-control-plane -it --profile=sysadmin --image=nicolaka/netshoot:v0.16 -- bash
Creating debugging pod node-debugger-ded-control-plane-sjpv8 with container debugger on node ded-control-plane.
All commands and output from this session will be recorded in container logs, including credentials and sensitive information passed through the command prompt.
If you don't see a command prompt, try pressing enter.
This Pod's spec shows that it shares the PID, network, and IPC namespaces with the node, runs privileged, and mounts the node's / at /host.
$ kubectl get pod/node-debugger-ded-control-plane-sjpv8 -o json | jq '{hostPID: .spec.hostPID, hostNetwork: .spec.hostNetwork, hostIPC: .spec.hostIPC, sc: .spec.containers[0].securityContext, mounts: .spec.containers[0].volumeMounts, volumes: .spec.volumes}'
{
"hostPID": true,
"hostNetwork": true,
"hostIPC": true,
"sc": {
"privileged": true
},
"mounts": [
{
"mountPath": "/host",
"name": "host-root"
},
...
"volumes": [
{
"hostPath": {
"path": "/",
"type": ""
},
"name": "host-root"
},
...
It's privileged, and on top of that it has the node's PID, network, and IPC namespaces and the node's filesystem. That makes it a Pod with very strong permissions. Every process on the node is visible, and so is the node's entire /. When you use it on a production node, keep in mind that this is the level of access you're working with.
Let's look under /host.
$ ls --color=never /host
LICENSES bin boot dev etc home kind lib lib64 media mnt opt proc root run sbin srv sys tmp usr var
This is the node's /.
Trace the cgroup from the app's process
The PID namespace is shared with the node, so you can find the app's process on the node with ps.
$ ps -eo pid,user,args | grep '[/]server -http='
4082 65532 /server -http=:3000
4118 65532 /server -http=:4000
:3000 is the app and :4000 is the sidecar, both running as uid 65532. Take the PID of the :3000 process.
$ P=$(pgrep -u 65532 -f '/server -http=:3000')
$ echo "P=$P"
P=4082
4082 is the same as app's PID on the node in the table from the background section. Reading /proc/<pid>/cgroup for this PID gives the path from the root of the node's cgroup hierarchy.
$ cat /proc/$P/cgroup
0::/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-pod719135d1_ee24_4f46_948a_440e1c6f4f77.slice/cri-containerd-698b44cb65ab91fba808118cba48f23ab4761ddd8a3e05977a04c4323bcef648.scope
The path is the part after 0::. Extract it.
$ CG=$(cut -d: -f3 /proc/$P/cgroup)
$ echo "CG=$CG"
CG=/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-pod719135d1_ee24_4f46_948a_440e1c6f4f77.slice/cri-containerd-698b44cb65ab91fba808118cba48f23ab4761ddd8a3e05977a04c4323bcef648.scope
Append this path to /host/sys/fs/cgroup, which corresponds to the node's /sys/fs/cgroup.
$ cat /host/sys/fs/cgroup$CG/cpu.max
50000 100000
50000 100000 in cpu.max is a limit of up to 50,000 microseconds of CPU time out of every 100,000 microseconds (100 ms). It matches the 500m limit set in "Test setup".
Whether the app actually hits this limit shows up in cpu.stat in the same directory.
$ cat /host/sys/fs/cgroup$CG/cpu.stat
usage_usec 56270
user_usec 34699
system_usec 21570
nice_usec 0
core_sched.force_idle_usec 0
nr_periods 22
nr_throttled 0
throttled_usec 0
nr_bursts 0
burst_usec 0
In cpu.stat, nr_periods is the number of 100 ms cpu.max periods in which the app tried to use the CPU. nr_throttled is the number of those periods in which it hit the limit and was throttled. In this run it was 0 out of 22, so the app isn't hitting the limit.
Limitations and caveats
Added containers can't be removed from the Pod
When you add a container with kubectl debug, its definition is appended to the Pod's spec. The process inside stops using resources once it exits, but the definition stays, even for containers that have already terminated. Here's the list after adding dbg-exit, which exits immediately because it runs -- true.
$ kubectl get pod app -o jsonpath='{.spec.ephemeralContainers[*].name}{"\n"}'
dbg-notarget dbg-default dbg-it dbg-legacy dbg-general dbg-baseline dbg-restricted dbg-netadmin dbg-sysadmin dbg-exit
Every container I added for this article's tests is on the list. An attempt to remove one is rejected by the API server.
$ kubectl patch pod app --subresource=ephemeralcontainers --type=json -p '[{"op":"remove","path":"/spec/ephemeralContainers/0"}]'
The Pod "app" is invalid: spec.ephemeralContainers: Forbidden: existing ephemeral containers "dbg-notarget" may not be removed
The record goes away only when the Pod is recreated. Until then the names stay taken as well. An attempt to add another container named dbg-exit was rejected with existing ephemeral containers "dbg-exit" may not be changed. If you experiment casually on a production Pod, these records stay there.
You can't use --target on a stopped container
An app that crashes right after starting and keeps restarting is in a state called CrashLoopBackOff. For testing, I prepared a Pod crash that exits right after it starts. It mounts the same ConfigMap as app at /config and gets the same environment variables (APP_MODE, APP_UPSTREAM).
$ kubectl get pod crash
NAME READY STATUS RESTARTS AGE
crash 0/1 CrashLoopBackOff 4 (81s ago) 2m50s
It's in CrashLoopBackOff and has already restarted 4 times. To find out why it crashed, start with the logs of the previous container.
$ kubectl logs crash -c app --previous
2026/09/27 02:31:58 exiting on purpose
exiting on purpose is the message from the deliberate exit I built in for this test.
If you add a debug container to this Pod with --target=app, it doesn't start (its state is CreateContainerError). The process it's supposed to share a PID namespace with isn't running.
The workaround that fits this constraint is the --copy-to option of kubectl debug. It creates a copy of the Pod and lets you replace a container in the copy and run it.
First, I replaced only the image with netshoot using --set-image.
$ kubectl debug crash --copy-to=crash-copy --set-image=app=nicolaka/netshoot:v0.16
$ kubectl get pod crash-copy -o jsonpath='{range .spec.containers[*]}{.name} {.image} {.args}{"\n"}{end}'
app nicolaka/netshoot:v0.16 ["-exit"]
The image is now netshoot, but the original Pod's argument -exit is still there.
$ kubectl get pod crash-copy
NAME READY STATUS RESTARTS AGE
crash-copy 0/1 RunContainerError 2 (4s ago) 20s
The copy didn't start either and ended up in RunContainerError. The reason is in the last termination state.
$ kubectl get pod crash-copy -o jsonpath='{.status.containerStatuses[0].lastState}{"\n"}'
{"terminated":{..."message":"failed to create containerd task: failed to create shim task: OCI runtime create failed: runc create failed: unable to start container process: error during container init: exec: \"-exit\": executable file not found in $PATH","reason":"StartError",...}}
The leftover argument -exit is treated as the command to run. It doesn't exist in netshoot, so the start fails.
So I specified the container to replace with --container=app, and passed both an image with --image and a command after --.
$ kubectl debug crash --copy-to=crash-copy2 --container=app --image=nicolaka/netshoot:v0.16 --image-pull-policy=Never -- sleep infinity
$ kubectl get pod crash-copy2 -o jsonpath='{range .spec.containers[*]}{.name} {.image} {.command} {.args}{"\n"}{end}'
app nicolaka/netshoot:v0.16 ["sleep","infinity"]
This time the command is sleep infinity and the arguments are empty.
$ kubectl get pod crash-copy2
NAME READY STATUS RESTARTS AGE
crash-copy2 1/1 Running 0 1s
The copy is Running and stays up.
In the copy's app container, the image and command are replaced, but the environment variables and volume mounts are carried over unchanged from the original Pod. So you can check the config file and environment variables the app would read at startup, under the same conditions.
$ kubectl exec crash-copy2 -c app -- cat /config/app.json
{"upstream": "http://backend.default.svc:8080", "timeout_ms": 250}
$ kubectl exec crash-copy2 -c app -- printenv APP_MODE APP_UPSTREAM
production
{"upstream": "http://backend.default.svc:8080", "timeout_ms": 250}
Both match what the original app receives. However, the copy doesn't run the app itself, so you can't see the process at the moment it crashes. What you can check is the set of conditions the app receives at startup.
Wrap-up
A distroless container has no shell, so you can't get in with kubectl exec. If you add an ephemeral container with kubectl debug, though, you can bring your own tools and investigate.
- The network is shared with the Pod, so
curl localhost,ss,tcpdump, anddigwork directly - If you enter with
--target, the app's process is visible, and you can read its state and environment variables from/proc/1 - The app's filesystem is separate and not directly visible. You can read it through
/proc/1/root -
/sys/fs/cgroupshows the debug container's own values, except with sysadmin. The app's values are readable through/proc/1/root. If an ephemeral container can't reach them, read them from the node
Whether you can read the contents of what you reach depends on two things.
- Inspecting another process requires
SYS_PTRACE. You get it by entering with general (no--profile) or sysadmin - If the target drops all capabilities (every target in this test did), reads worked without
SYS_PTRACEwhen your uid and gid both matched the target's
When you use this, watch out for two things.
- Added containers can't be removed from the Pod
- You can't use
--targeton a stopped container. For a Pod that keeps crashing, make a copy with--copy-toand investigate the copy
How far you can reach into the app depends on which namespaces you share with it. Whether you can read what you reach depends mainly on SYS_PTRACE. If the target drops all capabilities, matching uid and gid also let you read it. If you keep these two questions apart, you'll know which one to suspect when you get stuck.
Top comments (0)