DEV Community

DASWU
DASWU

Posted on

Nydus + JuiceFS: Reducing Container Startup Time for AI Inference from 116s to 1.4s

In large-scale AI inference services, when a cluster scales out, newly added inference instances must go through a series of cold-start steps before they can serve requests: container image preparation, file system mounting, runtime and inference framework initialization, and model weight loading. As image sizes and model weights continue to grow, the time spent on data preparation becomes increasingly prominent. During large-scale scaling events, many instances start simultaneously, further increasing the concurrent pressure on the registry, network, and backend storage.

To address these issues, the industry has developed several optimization strategies, including reducing image size, using P2P distribution to alleviate concurrent pull pressure, and employing on-demand loading to reduce data transfer on the critical startup path.

In this article, we’ll explore a shared‑storage‑based approach: combining Nydus’ on‑demand image loading capabilities with JuiceFS’ distributed storage and caching mechanisms. By sharing storage and reusing cached data, this approach reduces redundant data reads during large-scale scale-outs and shortens container startup time. For AI inference scenarios, container images and model weights can be managed together under a unified storage, caching, and prefetching system, accelerating both types of data while reducing the complexity of building and maintaining two separate data pipelines.

Why OCI image distribution is slow

Traditional Open Container Initiative (OCI) images are typically organized as compressed tar layers. In the image‑pull mode, a node must fully download and extract the required layers before the container can start. In practice, however, only about 6%–20% of files in an image are actually accessed during container startup; the rest of the download and extraction overhead is wasted on non‑critical‑path data.

In large clusters, this “prepare everything first, then start” approach causes three problems:

  • Long cold‑start times: For GB‑sized images (like AI inference images), full download and extraction significantly delay new instance startup, often taking minutes.
  • Registry bottleneck: During HPA‑triggered bursts or large‑scale rollouts, hundreds of nodes simultaneously request GB‑sized layers from the registry, easily overwhelming the registry and congesting the network.
  • Redundant network and storage usage: Without an efficient P2P or distributed sharing mechanism between nodes, the same layers are repeatedly transferred across the cluster, resulting in significant bandwidth consumption and duplicated storage.

Dragonfly+Nydus: what it solves and what it leaves behind

To mitigate registry bandwidth pressure during concurrent image pulls in large clusters, a common solution is to introduce P2P distribution systems like Dragonfly. The core idea is to build a P2P distribution network via a scheduler and node‑side agents, allowing nodes to share image data and reduce egress traffic from the registry.

However, P2P distribution primarily optimizes the data transfer path; it does not change the fundamental OCI mechanism of downloading and extracting layers before starting the container. Nydus, on the other hand, changes the loading approach: it transforms “prepare everything then start” into “fetch lightweight metadata first, and then read actual data on demand,” reducing data transfer on the critical startup path.

Combined, Nydus handles on‑demand loading, while Dragonfly distributes the on‑demand data, reducing both the actual image read volume and the concurrent access pressure on the registry.

However, this combination introduces additional infrastructure and operational complexity. Besides Nydus snapshotter, one must deploy and maintain Dragonfly’s scheduler, distributor, and node‑side components, and handle P2P network connectivity, cache space management, node anomalies, and cache eviction. In large‑scale GPU inference clusters, this means maintaining a separate data distribution and caching system independent of the storage layer.

This leads to an alternative thought: if the cluster already has a mature distributed file system like JuiceFS, could image data distribution and caching be offloaded to the shared storage system, reusing existing storage infrastructure while reducing the cost of building and maintaining extra P2P components?

Architecture redesign: Nydus blobs hosted by JuiceFS

The core idea of the Nydus+JuiceFS solution is to transform image data distribution—originally relying on P2P network scheduling—into shared storage and cache reuse based on a distributed file system. In this architecture, JuiceFS serves as the underlying unified data engine hosting Nydus image blob data, and worker nodes read data through the standard POSIX interface.

After an image is built and converted to Nydus format during the CI/CD phase, it’s split into two parts:

  • Bootstrap (metadata): Very small (typically a few hundred KB, for example, ~450KB), stored in the Image Registry, used for fast indexing.
  • Blobs (actual data blocks): Contains all file content (typically several GB, for example, ~1GB), persisted directly to the JuiceFS shared file system after conversion (for example, /var/lib/nydus/blobs).

When a worker node starts a container, containerd triggers nydus-snapshotter, and nydusd mounts the container’s read‑only layer via FUSE. Then, when the application issues actual read requests, nydusd looks up the index in the Bootstrap and reads the corresponding blob data on‑demand through the JuiceFS mount point.

The key change in this path is: the registry only distributes the manifest and the extremely lightweight bootstrap; the massive blob data transfer is entirely handled by JuiceFS’ distributed caching system. This fundamentally eliminates the registry network bottleneck during high‑concurrency pulls.

For AI inference scenarios, this architecture can further unify the storage paths for images and model weights. In addition to Nydus image blobs, model weights can also be stored in JuiceFS, allowing both types of data to reuse the same storage, caching, and prefetching mechanisms. This accelerates data loading while reducing the complexity of building and maintaining two data pipelines.

Comparison basis OCI full image pull Dragonfly+Nydus JuiceFS+Nydus
Core concept Full layer pull + layer extraction P2P network offloading + on‑demand image loading Storage‑layer decoupling + distributed caching + on‑demand loading
Cold start method & time Full download and extraction required; typically 30s to 10min+ Only download hundreds of kilobytes of bootstrap metadata and blobs needed for startup; usually seconds Only download hundreds of kilobytes of bootstrap metadata and blobs needed for startup; usually seconds
Image data transfer path Registry → worker (full local download & extraction) Registry / source → Dragonfly P2P network → worker on‑demand read JuiceFS → worker on‑demand read via POSIX, accelerated by local & distributed caches
Core components Registry + containerd Nydus, Manager, Scheduler, Seed Peer, and dfdaemon JuiceFS + nydus‑snapshotter
Image & model weight synergy Image pulled from the registry to local node; model weights loaded via a separate path (NAS or object storage); two systems managed independently Images loaded on‑demand via Nydus, distributed by Dragonfly; model weights still rely on independent external storage Image blobs and model weights stored together in JuiceFS, sharing the same cache and warm-up mechanisms
Operational complexity Easy to overwhelm registry bandwidth under high concurrency Requires extra maintenance of P2P scheduler, node network, and caching system Can reuse existing JuiceFS storage & cache infrastructure, reducing independent image distribution components
Local disk pressure High: Stores full image and extracted data High and unpredictable: Requires cache space for P2P; rapid scaling may cause disk full Controllable: Unified shared storage pool, uses local NVMe for P2P distributed cache
Network requirements Standard TCP Requires widely open P2P ports across nodes; may be invalid in cross‑VPC/NAT environments Standard LAN – only requires normal high‑bandwidth inter‑node LAN; no intrusive network requirements
Suitable scenarios General workloads with no cold‑start requirements and small clusters (<20 nodes) Ultra‑large clusters across data centers / clouds / edge computing environments without unified high‑bandwidth storage; strong ops capabilities of the team Ultra-large clusters across data centers / clouds / edge computing environments without unified high-bandwidth storage, where JuiceFS is already deployed or can be deployed, aiming for unified image+model management and simplified ops

Runtime mechanism: how Nydus and JuiceFS coordinate

Startup phase: fetching bootstrap (metadata)

When the application starts, nydusd only needs to fetch the lightweight bootstrap file tree from the registry. The bootstrap describes the image file tree and block indexes, but contains no actual file content.

After the container root filesystem (rootfs) is mounted, the subsequent startup process proceeds. In our test scenarios, containers can be ready and pass health checks within hundreds of milliseconds, without transferring the full large image data before startup.

Read phase: on‑demand blob access

When the application process actually issues read operations (such as open() or read()), nydusd captures the I/O requests via the Linux FUSE mechanism and locates the corresponding blob data block and offset using the bootstrap index.

This process is completely transparent to upper‑layer applications. For the underlying data path, the read granularity changes from full image layers to the data blocks corresponding to specific files and offsets. nydusd directly accesses the JuiceFS mount point (for example, /var/lib/nydus/blobs) to read the target block, completely eliminating the runtime fetch‑from‑registry path.

Cache phase: data reuse via JuiceFS cache groups

JuiceFS’ distributed cache groups can reuse already‑cached blob data across the cluster. On a cache hit, data is returned to the kernel at near‑local‑NVMe speed. On a cache miss, only the specified offset block is pulled on‑demand and automatically filled into the cache. In addition to the automatic fill‑on‑miss, blobs can be pre-loaded into the cache to ensure full cache hits during on‑demand reads.

For AI inference scenarios, the same mechanism can also be applied to loading model weight files.

Production test data

In production, we compared the solutions above against a large AI inference image.

Container startup time comparison

The table below shows the test results for container startup time:

Scenario OCI overlayfs Nydus+JuiceFS Comment
Cold pull + run 116 seconds 16 seconds OCI uses public cloud registry; Nydus+JuiceFS uses JuiceFS + public cloud object storage
Cold pull + run (with distributed cache warm-up) N/A ~1.4 seconds Required blobs pre‑loaded in JuiceFS distributed cache
Warm pull + run ~0 seconds 0.22 seconds OCI: same tag already on node; Nydus+JuiceFS: blobcache local cache

From the test results, we can see:

  • 7x speedup without cache warm-up: When neither local nor distributed cache is warmed up, Nydus+JuiceFS reduces container startup time from 116s to 16s.
  • ~99.9% reduction in data read volume during startup: In the cold pull + run scenario, read volume drops from 11 GB to 11 MB. Traditional OCI requires full download and extraction before startup, while Nydus+JuiceFS only fetches bootstrap metadata and reads blobs actually accessed during startup.
  • 1.4s startup with distributed cache warm-up: When blobs are cached in JuiceFS distributed cache, startup takes 1.4s; with local blobcache warmed up (warm pull + run), startup takes only 0.22s.

On‑demand file read performance

The table below shows test results for: on-demand file read performance:

Scenario Traditional OCI (local disk direct read) Nydus+JuiceFS (distributed cache hit)
First‑block read latency (TTFB) ~0.9 milliseconds ~3.2 milliseconds
Small file / scattered configuration loading (IOPS) Baseline Close to baseline
Large file sequential read throughput Local disk limit GB/s level (JuiceFS NVMe distributed cache)

From the test results, we can see:

  • Time to first byte (TTFB): On distributed cache hit, first‑block read latency is ~3.2 milliseconds, about 2.3 milliseconds higher than local disk direct read (0.9 milliseconds). The extra latency comes from the FUSE data path and cache access, but remains at millisecond level, generally imperceptible to the application side.
  • Small file IOPS: Performance is very close to local disk baseline. In scenarios with dense Python dependency and configuration file reads, JuiceFS caching introduces no noticeable IOPS bottleneck.
  • Large file sequential throughput: On cache hit, Nydus+JuiceFS achieves GB/s sequential read throughput. Compared to traditional OCI, which is limited by single‑node local disk performance, JuiceFS can leverage NVMe distributed cache to provide high‑bandwidth data access.

Summary: combining on‑demand loading with shared cache

Nydus changes image loading from full pull to on‑demand reading, while JuiceFS places data blocks into a shared storage and distributed cache system. Together, they reduce container cold‑start time from 116s to 16s, and to ~1.4s with distributed cache warm-up; the data read volume during startup drops from 11 GB to 11 MB.

The value of this solution extends beyond just image startup acceleration. For AI inference scenarios, image blobs and model weights can be uniformly hosted by JuiceFS, sharing the same caching and prefetching mechanisms. During large‑scale scaling, distributed caching also reduces duplicate reads of the same data across many instances. This lessens concurrent pressure on the registry and backend storage.

In environments lacking shared storage, Dragonfly+Nydus is an excellent P2P alternative. However, in clusters that already have JuiceFS infrastructure, Nydus+JuiceFS delivers extreme image distribution and model loading acceleration with minimal operational overhead.

If you have any feedback on this article or ideas to share, we invite you to participate in the discussions on GitHub and join our community on Discord.

Top comments (0)