A checkpoint restore can rebuild a process from saved state without passing that state back through the destination's normal policy translation. When it does, the security context on the Pod spec describes the workload that was approved, and the process on the node is the one that was saved.
That is a boundary problem before it is a vulnerability problem. Kubernetes has one place where policy becomes process state, and it is creation. Restore is a second route to a running process, and nothing on that route decides what wins when saved state and destination policy disagree. Advisories on the restore path in containerd and CRI-O show what crosses when nothing decides. They are evidence for the architecture, and the architecture is the subject.
Two Ways to Create a Process
On the normal path, the process is built from declarations. A Pod spec passes admission, the kubelet hands the runtime a container config, and the runtime turns that config into kernel state: a user and group, a capability set, the no_new_privs flag, a seccomp filter. That translation step is how modern infrastructure architecture turns declared intent into enforced state. Each attribute is policy made concrete, which is why seccomp filters and capability drops eliminate whole classes of exploit only when they are actually attached to the process. Kubernetes documents that allowPrivilegeEscalation directly controls whether no_new_privs is set. The field in the spec is the request. The flag on the running process is the enforcement.
Checkpoint restore does not build a process. It reconstitutes one. CRIU records a running process, including its credentials, capabilities, no_new_privs flag and seccomp state, and restore replays that record. On the path containerd exposed, the restore is triggered inside the ordinary container-create call, by a checkpoint archive or an annotated image. Kubernetes currently supports container restore only through those image annotations, so from admission's side this is a Pod create with an image reference. Whether it is a restore is decided later, inside the runtime, by what the image contains.
| Stage | Create path | Restore path |
|---|---|---|
| Source of process state | Pod spec, defaulted and admitted | Saved process image |
| Policy translation | Runtime turns the container config into uid, capabilities, no_new_privs, seccomp |
Not applied to the restored attributes |
| What admission sees | A Pod with a spec | A Pod with an image reference |
| Who decides it is a restore | Not applicable | The runtime, from the image contents |
Restore is usually treated as a data operation whose failures show up in the layers above it, the pattern restore design failure documents: the restore completes and the workload is still unusable. This is a different gap. Every layer can pass, and the process can still come back carrying state the destination never approved.
What the Advisories Show About Checkpoint Restore
containerd's September 1 advisory states the outcome directly. A container restored from an untrusted checkpoint through the create call can run as root with full capabilities and no seccomp filter, despite restrictive policy requested by the orchestrator. It applies where checkpoint restore through CRI is enabled and an attacker can run a container from a crafted checkpoint image. The affected ranges run from 2.1.0 up to, but not including, 2.2.7, and from 2.3.0 up to, but not including, 2.3.4. Those two releases disable the path by default.
The same checkpoint and restore path had already produced three other trust failures in June, according to Google's GKE security bulletins. Each accepted something the checkpoint supplied.
| What the artifact was trusted for | Advisory | What restore accepted without translation |
|---|---|---|
| Image identity | CVE-2026-50195, June | Image references in a checkpoint import, unvalidated, allowing a poisoned node image cache |
| Devices and host mounts | CVE-2026-53492, June | CDI annotations carried in checkpoint metadata, bypassing resource allocation and device plugin enforcement |
| Host file reads | CVE-2026-53489, June | A symlinked log path restored without validation, allowing arbitrary host file reads through kubectl logs |
| Process privilege | GHSA-p7v4-vr35-mj6f, September | Credentials, capabilities, no_new_privs and seccomp state from the checkpoint |
CRI-O's CVE-2026-92574 records the same outcome for that runtime: saved credentials, capabilities, no_new_privs and seccomp state can take effect where the destination's configuration should have.
⚠ Scope of exposure: Each of these needs the ability to create Pods, and the September issue matters only where checkpoint restore is enabled. containerd rates it Critical. Google rates it Medium for GKE and notes that default GKE nodes ship without the criu binary. The point here is the boundary, not the count of exposed clusters.
Precedence Was Never Decided
These are not four bugs in four attributes. Each is the same event: state supplied by an artifact took effect on restore, and the destination's policy was not applied over it. Nothing adjudicated. The runtime did not weigh the saved credentials against the requested user and pick the saved ones. The translation that would have replaced them never ran. That difference matters. A precedence rule that was chosen can be audited and changed. One that was never made has to be built.
Who has the right to change infrastructure is the question the Control Plane Boundaries stage of the Modern Infrastructure learning path puts at the center, with GitOps, CI/CD, consoles and platform APIs as competing claimants. A restore path is one more claimant, and it arrives with no place in that ordering. It is also a different condition from Policy Intent Drift (#133), where declared and enforced state match and the reason for the rule has expired. On restore, declared and effective state never matched, and nothing compared them.
containerd's handling shows the shape of it. The June issues were fixed in point releases, the CDI one in 2.1.9, 2.2.5 and 2.3.2, with the restore path left in place. The September issue could not be fixed that way: in the advisory's own terms, containerd "cannot enforce destination security policy" while it restores the process. So 2.2.7 and 2.3.4 disable the path by default, and 2.4 removes it. That removal was already scheduled. containerd's release documentation lists restore during CRI create as deprecated in 2.3, released April 30, with removal targeted for 2.4 and the RestorePod API in KEP-5823 named as the replacement. The advisories did not set that timeline. They show why the path cannot be kept safe by patching one attribute at a time.
Status Is Not State
containerd's advisory adds a second problem. Its CRI status reporting reflects the requested configuration, not the state of the restored process, so an orchestrator reading status sees the destination policy while the process runs the saved one. Google's bulletin repeats the point. The CRI-O record does not describe status reporting, so the claim belongs to containerd alone.
This is the split Kubernetes requests and limits already draws for CPU and memory: the spec records what was requested, and the kernel enforces what the process holds. The difference here is that the checkpoint restore path can make the two disagree on security attributes while the control plane keeps reporting the request. The evidence has to come from the process. On a node, the NoNewPrivs, Seccomp and CapEff fields in /proc/<pid>/status are what to compare against the security context the Pod asked for.
grep -E 'NoNewPrivs|Seccomp|CapEff' /proc/<pid>/status
These fields reveal the effective process state after restore and can be compared against the Pod's requested security context.
Where Support Windows Make This Live
Whether the mitigation is available depends on the release line. containerd's release documentation lists 2.1 as end of life on July 3, 2026. The September advisory's affected range starts at 2.1.0, and its patched versions are 2.2.7 and 2.3.4, so it lists no fixed 2.1 release. On unpatched versions the advisory says containerd offers no option to disable restore through the create call. Exposure still requires checkpoint restore to be in use, but a support date decides whether turning it off is a configuration change or an upgrade. A support date and an enforcement date are different things, as the Kubernetes 1.37 containerd deadline shows for a different deadline. Version 2.2 is active until November 6, 2026, 2.3 is the LTS release through April 2028, and 2.4 has removed the path.
What Restore Has to Prove Before It Runs
If restore is a second instantiation path, it needs its own proof, because admission cannot supply it. Three things have to be shown. That a restore was requested, deliberately, by someone allowed to request it. That the artifact came from a source the platform trusts. And that the process that came back holds the attributes the destination policy specifies. Google's bulletin turns the first two into controls: restrict who can create Pods, allow only verified image registries, and watch node logs for restore entries and the runtime's events for the deprecation warning on restore through create. The third is the one status cannot give you.
The replacement makes the first requirement explicit. Under KEP-5823, which targets alpha in Kubernetes 1.38 and has not shipped, restore is declared on the Pod spec, authorized through a dedicated restore verb, and admitted only if the Pod spec equals a template the kubelet recorded when the checkpoint was taken. That settles the conflict by forbidding divergence between spec and checkpoint. It does not, as the KEP is written, verify the restored process itself: the equality check runs against the recorded object, and the runtime's checkpoint archive is treated as opaque to Kubernetes. The KEP's text does not describe checking restored credentials, capabilities, no_new_privs or seccomp state against the spec.
A restore that cannot enforce destination policy should not run. That is the position containerd reached when it disabled the path, and it is the rule any platform allowing restore has to state for itself: who decides what a restored workload may hold, and what is read from the process to prove it.
Diagnostic: "When a Pod is created from an image that contains checkpoint data, which component decides its security context, and what would you read to find out?"
Architect's Verdict
Normal creation applies destination policy. Checkpoint restore reconstructs authority from saved state without that translation, and nothing decides which of the two wins.
The instinct is to read this as a run of containerd bugs, each patched as it surfaced. The bugs are what crossed the boundary. The boundary is that restore is a separate instantiation path with no owner for precedence, and a runtime that cannot enforce destination policy while it restores has no safe way to keep the path open.
Admission approves a spec. It cannot approve a process it never sees constructed. Any platform that allows checkpoint restore has to say who decides what a restored workload may hold, and prove it by reading the process, not the spec.
Admission decides what a workload may be. Restore decides what it becomes when nobody asks.
Download: The Checkpoint Restore Policy Gap Carousel (PDF, 10 slides)
Originally published at rack2cloud.com



Top comments (0)