<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: DASWU</title>
    <description>The latest articles on DEV Community by DASWU (@daswu).</description>
    <link>https://dev.to/daswu</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F973459%2F8329f19d-864e-41c4-bb30-a5a0d23e63ce.jpeg</url>
      <title>DEV Community: DASWU</title>
      <link>https://dev.to/daswu</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/daswu"/>
    <language>en</language>
    <item>
      <title>JuiceFS Enterprise 5.4: From 100 Billion Files to 1 Million Clients</title>
      <dc:creator>DASWU</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:21:31 +0000</pubDate>
      <link>https://dev.to/daswu/juicefs-enterprise-54-from-100-billion-files-to-1-million-clients-1760</link>
      <guid>https://dev.to/daswu/juicefs-enterprise-54-from-100-billion-files-to-1-million-clients-1760</guid>
      <description>&lt;p&gt;Following the support for 100-billion-file environments introduced in &lt;a href="https://juicefs.com/docs/cloud/" rel="noopener noreferrer"&gt;JuiceFS Enterprise Edition&lt;/a&gt; 5.3, v5.4 further improves JuiceFS across several dimensions for large-scale deployments. At this scale, even small per-operation overheads can add up to significant resource consumption. Once metadata is distributed across multiple zones, the system also needs to coordinate across nodes while maintaining performance, consistency, and stability.&lt;/p&gt;

&lt;p&gt;To address these challenges, &lt;a href="https://juicefs.com/docs/cloud/release#juicefs-543-2026922" rel="noopener noreferrer"&gt;JuiceFS Enterprise Edition 5.4&lt;/a&gt; introduces the following new capabilities and improvements:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Support for up to 1 million simultaneously mounted clients
&lt;/li&gt;
&lt;li&gt;Fast cloning of directories containing massive files
&lt;/li&gt;
&lt;li&gt;More than 90% RDMA bandwidth utilization
&lt;/li&gt;
&lt;li&gt;Up to 100% higher random-read performance under high concurrency
&lt;/li&gt;
&lt;li&gt;On-demand data synchronization for mirror file systems
&lt;/li&gt;
&lt;li&gt;More flexible caching strategies for a wider range of workloads
&lt;/li&gt;
&lt;li&gt;Multiple operational improvements for large-scale data&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Supporting 1 million clients
&lt;/h2&gt;

&lt;p&gt;As &lt;a href="https://en.wikipedia.org/wiki/AI_agent" rel="noopener noreferrer"&gt;AI agents&lt;/a&gt; move toward large-scale deployment, the JuiceFS team faced a new challenge: &lt;strong&gt;supporting connections and concurrent access from up to 1 million clients on top of the metadata workload of a 100-billion-file file system.&lt;/strong&gt; The team spent months exploring and optimizing the architecture, and introduced this capability in JuiceFS 5.4.&lt;/p&gt;

&lt;p&gt;JuiceFS uses a &lt;a href="https://juicefs.com/docs/cloud/introduction/architecture" rel="noopener noreferrer"&gt;distributed architecture&lt;/a&gt; that partitions metadata across multiple metadata nodes, allowing metadata for hundreds of billions of files to be managed across multiple zones. To allow these nodes to handle a large number of client connections at the same time, v5.4 co-locates metadata nodes with network proxies, enabling clients to connect to any metadata node.&lt;/p&gt;

&lt;p&gt;The node that receives a request handles requests for its local partition directly. Requests involving other partitions are forwarded internally to the corresponding nodes. Multiple client sessions share the same physical connections between nodes, reducing the number of connections and the associated maintenance overhead.&lt;/p&gt;

&lt;p&gt;Connection reuse is only part of the challenge. Querying large numbers of client sessions, reporting session status, and cleaning up expired sessions can also introduce significant overhead.&lt;/p&gt;

&lt;p&gt;In v5.4, indexes are used to locate client sessions instead of scanning unrelated sessions. Paginated queries and sampled reporting reduce the overhead of session status management. Expired sessions are cleaned up in batches to prevent cleanup work from becoming concentrated in a single pass.&lt;/p&gt;

&lt;p&gt;With these optimizations, JuiceFS 5.4 successfully handled 1 million simultaneous client connections in stress tests. Previously, when clients connected separately to multiple metadata nodes, the total number of connections could reach tens of millions. &lt;strong&gt;After optimization, 10 internal proxy nodes handled 1 million client connections, while connection reuse reduced the number of physical backend connections from the proxies to the metadata leader to about 400. This significantly reduced connection-management pressure on the leader.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffuffqdezigzmkmdfkcev.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffuffqdezigzmkmdfkcev.png" alt=" " width="800" height="267"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Fast cloning for large directories
&lt;/h2&gt;

&lt;p&gt;When creating branches of training data or preparing test environments, users often need an independently modifiable copy of a directory. As directory sizes grow, &lt;a href="https://juicefs.com/docs/cloud/guide/clone/" rel="noopener noreferrer"&gt;cloning&lt;/a&gt; becomes more complex: &lt;strong&gt;in a multi-zone architecture, the metadata for a single directory tree may span multiple metadata zones, so cloning must coordinate subtree operations across nodes while preserving the complete directory hierarchy.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;v5.4 introduces cross-zone directory cloning. It recursively processes subtrees while preserving the original cross-zone layout of the directory tree. Cloning only copies metadata. The underlying object data remains shared, so there is no need to copy it. Subsequent writes to the clone do not affect the source files.&lt;/p&gt;

&lt;p&gt;In a deployment with 30 metadata zones, cloning 100 million files took 1 minute and 40 seconds, &lt;strong&gt;averaging 1 million files per second&lt;/strong&gt;. More zones can provide greater parallelism, while actual performance also depends on metadata load and how the directory tree is distributed across zones.&lt;/p&gt;

&lt;p&gt;Note that cross-zone directory cloning is not atomic. If new data continues to be written to the source directory during cloning, different subtrees in the cloned directory may reflect the source directory at different points in time. If strict consistency is required, pause writes to the source directory before starting the clone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Performance improvements
&lt;/h2&gt;

&lt;h3&gt;
  
  
  More than 90% RDMA bandwidth utilization
&lt;/h3&gt;

&lt;p&gt;Starting with JuiceFS 5.3, &lt;a href="https://juicefs.com/en/blog/release-notes/juicefs-enterprise-5-3-rdma-support#Support-for-RDMA:-increased-bandwidth-cap,-reduced-CPU-usage" rel="noopener noreferrer"&gt;clients can use RDMA&lt;/a&gt; to communicate with distributed cache nodes, reducing CPU overhead during data transfers and improving transfer efficiency. v5.4 further improves bandwidth utilization and transfer stability.&lt;/p&gt;

&lt;p&gt;In a test environment with a 400 Gbps RDMA NIC on the client and an 800 Gbps RDMA NIC on the distributed cache node, a single client achieved 45 GB/s of read throughput. A single distributed cache node delivered 90 GB/s of data throughput, with NIC utilization exceeding 90%.&lt;/p&gt;

&lt;h3&gt;
  
  
  Higher random-read IOPS under high concurrency
&lt;/h3&gt;

&lt;p&gt;In multi-process training, HPC, and rendering scenarios, high-concurrency random reads can stress not only storage and network resources, but also client lock contention and request-processing overhead.&lt;/p&gt;

&lt;p&gt;Over the past six months, the JuiceFS team has continuously optimized high-concurrency random-read performance by simplifying frequently executed code paths and reducing lock contention.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;As a result, random-read IOPS has more than doubled overall compared with the previous implementation.&lt;/strong&gt; With distributed caching, a single client reached up to 189,000 IOPS. With local caching, a single client reached up to 156,000 IOPS.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs2lglokqwnsbpdniod2d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs2lglokqwnsbpdniod2d.png" alt=" " width="800" height="361"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Test environment:&lt;/strong&gt; A dual-socket AMD server with 128 cores and 256 threads, 1 TB of DDR4 memory, a 100 Gbps NIC, and a 3.84 TB NVMe SSD used as the cache disk. The system ran Linux 5.14.&lt;/p&gt;

&lt;h2&gt;
  
  
  Optimizing data access across regions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  On-demand synchronization to reduce data replication
&lt;/h3&gt;

&lt;p&gt;In multi-region, multi-provider, and multi-cluster compute environments, a JuiceFS &lt;a href="https://juicefs.com/docs/cloud/guide/mirror/" rel="noopener noreferrer"&gt;mirror file system&lt;/a&gt; can automatically synchronize files and directories from the source. Metadata is continuously synchronized to keep the directory structure and object versions up to date. Previously, object data was synchronized in full in the background by default.&lt;/p&gt;

&lt;p&gt;However, a workload in the mirror region may only need a subset of the models, datasets, or historical directories stored at the source. For example, if the source contains 10 PB of data but the mirror accesses only 1% of it, full synchronization would transfer and store a large amount of data that the workload never uses.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;v5.4 supports on-demand synchronization of object data for mirror file systems. When enabled, metadata continues to synchronize, while object data is no longer synchronized automatically in the background.&lt;/strong&gt; When a client reads data, JuiceFS first checks the object storage in the mirror region. If the data has not been synchronized yet, JuiceFS fetches it from the source and stores it in the mirror on demand. This reduces unnecessary data transfer and storage consumption.&lt;/p&gt;

&lt;p&gt;If an application requires local read performance on the first access, you can warm up the required data in advance. If a complete copy of the data is required, you can continue to use background synchronization.&lt;/p&gt;

&lt;h3&gt;
  
  
  Accelerating metadata access with read-only nodes
&lt;/h3&gt;

&lt;p&gt;For workloads that only need to accelerate metadata access, &lt;strong&gt;v5.4 introduces a new mode that does not require creating a mirror file system&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Users can deploy read-only metadata service nodes in other regions to accelerate metadata reads for local clients. Clients continue to mount the source file system, while the acceleration nodes can scale dynamically to match access requirements.&lt;/p&gt;

&lt;h2&gt;
  
  
  More flexible cache management
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Improving hot-data cache hit rates
&lt;/h3&gt;

&lt;p&gt;Data from one-time scans is often unlikely to be accessed again. Caching this data can evict hot data that is accessed repeatedly. By default, when a request misses the distributed cache, the cache node fetches the data from object storage and stores it in the cache. &lt;strong&gt;v5.4 introduces a new cache-aside mode.&lt;/strong&gt; When enabled, a cache miss causes the application node to read directly from object storage instead of filling the cache with the requested data.&lt;/p&gt;

&lt;h3&gt;
  
  
  Preserving cache hit rates when adding multiple cache nodes
&lt;/h3&gt;

&lt;p&gt;When multiple cache nodes are added at once, read requests may be routed to new nodes that do not yet have the required data, triggering additional reads from object storage.&lt;/p&gt;

&lt;p&gt;v5.4 introduces the &lt;code&gt;--commission&lt;/code&gt; parameter. After a new node is enabled, it retains the node mapping from before scaling. If the data is not available in its local cache, the new node fetches it from the original node instead. This reduces reads from the source during cache scaling.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reducing the number of files on large cache disks
&lt;/h3&gt;

&lt;p&gt;On large-capacity cache disks, storing each cache block as a separate file can cause the number of files on the cache disk to grow continuously.&lt;/p&gt;

&lt;p&gt;The new merge-cache mode combines multiple cache blocks into larger files and stores their indexes separately. This significantly reduces the number of files that the cache disk needs to manage.&lt;/p&gt;

&lt;p&gt;This feature is currently available as a public beta.&lt;/p&gt;

&lt;h3&gt;
  
  
  Warming up specific byte ranges
&lt;/h3&gt;

&lt;p&gt;If a workload reads only part of a large file, warming up the entire file consumes additional cache space and transfers unnecessary data.&lt;/p&gt;

&lt;p&gt;v5.4 allows you to specify one or more byte ranges to warm up. JuiceFS only warms up cache blocks that intersect those ranges, aligning cache preparation with the data the workload actually needs to read.&lt;/p&gt;

&lt;p&gt;This capability enables more precise warm-up for data formats such as Lance and Parquet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Simplifying operations for massive files
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Restoring files by deletion time
&lt;/h3&gt;

&lt;p&gt;When restoring files from the trash, users may know approximately when an accidental deletion occurred but have difficulty identifying the target files by name or keyword alone.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;restore&lt;/code&gt; command now supports the &lt;code&gt;start-time&lt;/code&gt; and &lt;code&gt;end-time&lt;/code&gt; parameters. In addition to the existing keyword filters, these parameters allow users to restrict the restore operation by file deletion time.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reducing GC and &lt;code&gt;fsck&lt;/code&gt; memory usage
&lt;/h3&gt;

&lt;p&gt;When checking large file systems, &lt;a href="https://en.wikipedia.org/wiki/Garbage_collection_computerscience" rel="noopener noreferrer"&gt;GC&lt;/a&gt; and &lt;code&gt;fsck&lt;/code&gt; need to process large numbers of records, and sorting can consume significant amounts of memory.&lt;/p&gt;

&lt;p&gt;v5.4 adds external sorting modes to both commands, using disk storage during sorting to reduce memory requirements. The external sorting mode for GC is available only with dry-run, allowing users to inspect and evaluate the operation without actually deleting data.&lt;/p&gt;

&lt;h3&gt;
  
  
  Restricting access with token allowlists
&lt;/h3&gt;

&lt;p&gt;v5.4 supports access control through token allowlists. An allowlist can specify multiple first-level subdirectories under the root directory.&lt;/p&gt;

&lt;p&gt;Clients using the token can only access and modify content within those subdirectories, allowing different teams or workloads to be restricted to their designated data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Summary: scale changes everything
&lt;/h2&gt;

&lt;p&gt;JuiceFS 5.3 brought the platform into a new phase of scale. At the same time, the rapid development of AI continues to push the boundaries of data volume and system load.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;As file counts, client counts, and request volumes all reach new levels, performance, stability, consistency, and operational complexity are no longer independent concerns. They interact with one another and become increasingly difficult to manage as the system scales.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;v5.4 addresses these system-level engineering challenges across large-scale deployments, covering client connectivity, cross-zone directories, data access, mirror synchronization, and cache management. JuiceFS Cloud Service users can now try &lt;a href="https://juicefs.com/docs/cloud/release#juicefs-543-2026922" rel="noopener noreferrer"&gt;JuiceFS Enterprise Edition 5.4&lt;/a&gt; online. Users running on-premises deployments can contact the JuiceFS team for upgrade support.&lt;/p&gt;

&lt;p&gt;JuiceFS is used across a range of demanding workloads, including &lt;a href="https://en.wikipedia.org/wiki/Large_language_model" rel="noopener noreferrer"&gt;LLMs&lt;/a&gt; and multimodal models, autonomous driving, embodied AI, and quantitative investment. We’ll continue working with organizations at the forefront of these fields to improve JuiceFS and build infrastructure that can keep pace with their applications.&lt;/p&gt;

&lt;p&gt;If you have any feedback on this article or ideas to share, we invite you to participate in the &lt;a href="https://github.com/juicedata/juicefs/discussions/" rel="noopener noreferrer"&gt;discussions on GitHub&lt;/a&gt; and join &lt;a href="http://go.juicefs.com/discord" rel="noopener noreferrer"&gt;our community on Discord&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Infinite Workspaces for AI Agents: A Practical Architecture with JuiceFS, SQLite, and Litestream</title>
      <dc:creator>DASWU</dc:creator>
      <pubDate>Thu, 24 Sep 2026 01:58:23 +0000</pubDate>
      <link>https://dev.to/daswu/infinite-workspaces-for-ai-agents-a-practical-architecture-with-juicefs-sqlite-and-litestream-lhp</link>
      <guid>https://dev.to/daswu/infinite-workspaces-for-ai-agents-a-practical-architecture-with-juicefs-sqlite-and-litestream-lhp</guid>
      <description>&lt;p&gt;Most conversations about AI sandboxes start and end with boot speed and compute isolation: How fast does the container boot? Can the &lt;a href="https://en.wikipedia.org/wiki/AI_agent" rel="noopener noreferrer"&gt;agent&lt;/a&gt; run untrusted code safely? But these discussions miss a fundamental truth: an autonomous agent turns a fresh machine into a rambling digital workspace. It clones code repos, installs heavy &lt;code&gt;node_modules&lt;/code&gt; or Python virtual environments, downloads browser binaries, compiles artifacts, generates multi-gigabyte datasets, and writes massive trace logs.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;"The files are not incidental. They are the working memory of the task."&lt;/em&gt;&lt;br&gt;&lt;br&gt;
&lt;em&gt;– Aniket Maurya, Founder @ Celesto AI&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The naive solution is to overprovision the root disk. However, determining the size of a boot disk before a task exposes its scope is an ineffective effort. A "quick bug fix" might require a full Next.js build chain, while a "data benchmark" might produce multi-gigabyte Parquet tables. Coupling system state to workspace state creates a bottleneck.&lt;/p&gt;

&lt;p&gt;Instead, an emerging architectural pattern is radical in its simplicity: decouple the system disk from the durable workspace. Provide a POSIX-compliant mount that acts like an infinite, persistent drive, backed entirely by affordable object storage. But how do you build such a system without introducing complex services that defeat the purpose of ephemeral sandboxes? Recently, we co-hosted &lt;a href="https://juicefs.com/en/events/juicefs-office-hours-13-petabyte-scale-storage-for-ai-agent-sandboxes" rel="noopener noreferrer"&gt;an Office Hours session&lt;/a&gt; with &lt;a href="https://celesto.ai/" rel="noopener noreferrer"&gt;Celesto AI&lt;/a&gt;, a platform building secure sandboxes for autonomous AI agents, and the webinar was a delightful exploration of exactly this challenge.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Storage&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Root disk&lt;/td&gt;
&lt;td&gt;Operating system, runtime, package manager internals, and system-level state&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Workspace&lt;/td&gt;
&lt;td&gt;Repositories, generated files, datasets, build artifacts, logs, and project state&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Here is the creative engineering stack making this possible: &lt;a href="https://juicefs.com/docs/community/introduction/" rel="noopener noreferrer"&gt;JuiceFS&lt;/a&gt; as the file system layer, SQLite as the embedded metadata engine, and Litestream as the durability glue.&lt;/p&gt;

&lt;h2&gt;
  
  
  The creative trio: JuiceFS + SQLite + Litestream
&lt;/h2&gt;

&lt;p&gt;Object stores (S3, GCS) are excellent for storing files and performing key-based lookups, but they are terrible at listing directories, handling &lt;code&gt;stat&lt;/code&gt; calls, or managing millions of small files in batches (i.e., renaming a directory). You need a metadata engine to act as the "brain" of the file system, mapping paths to block locations.&lt;/p&gt;

&lt;p&gt;Usually, JuiceFS deployments use centralized or clustered metadata services, such as Redis, TiKV, or FoundationDB. For isolated AI sandboxes, a creative alternative leverages the three familiar JuiceFS components working in harmony:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;The &lt;strong&gt;JuiceFS client&lt;/strong&gt; provides the POSIX bridge. It is responsible for the complex task of communicating with both the metadata engine and the object store. It is the "worker" that reads and writes the actual data, as well as performs administrative tasks such as status checks and data compaction.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;SQLite&lt;/strong&gt; acts as the metadata engine. Instead of connecting to a remote metadata server, JuiceFS also supports using a local SQLite database file as its metadata store. This means the entire file system metadata (every inode, directory entry, and block pointer) lives inside a single, standard &lt;code&gt;.db&lt;/code&gt; file on the local disk.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Litestream&lt;/strong&gt; provides the durability backing for that SQLite file. Litestream continuously takes SQLite snapshots and streams the SQLite write-ahead log (WAL) frames to the same object store.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Here is the magic: When the sandbox is running, JuiceFS uses the local SQLite file to resolve &lt;code&gt;ls&lt;/code&gt;, &lt;code&gt;cd&lt;/code&gt;, and &lt;code&gt;read&lt;/code&gt; operations. The actual file blocks are already safe in the object store. (Although, the &lt;a href="https://litestream.io/reference/config/#replica-settings" rel="noopener noreferrer"&gt;&lt;code&gt;sync-interval&lt;/code&gt; configuration&lt;/a&gt; needs to be taken into consideration.) When the user stops the sandbox, the local SQLite file might be discarded, but Litestream has already replicated the metadata engine to the object store.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhtzd3wb83rf092067bpv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhtzd3wb83rf092067bpv.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When the sandbox resumes on a fresh VM (or a different host entirely), the system simply downloads the latest SQLite snapshot from the object store and replays the WAL. Within seconds, the local SQLite can be fully reconstructed. JuiceFS mounts it, points back to the same object store blocks, and the agent is looking at the exact same workspace directory it left behind.&lt;/p&gt;

&lt;p&gt;The JuiceFS file system doesn't care which computer it runs on. But it does remember its folder structure and file pointers with the portable SQLite snapshots and WAL frames. With this setup, we turned the working memory of the AI agent into files and blocks in the object store, making it durable and accessible instead of locking it to a specific boot disk or block device.&lt;/p&gt;

&lt;h2&gt;
  
  
  An unconventional but effective combo
&lt;/h2&gt;

&lt;p&gt;JuiceFS supports interchangeable metadata engines. Choosing SQLite + Litestream subverts the typical playbook by prioritizing simplicity and per-agent isolation over global scale. Here is why this particular stack fits the AI agent use case so well:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Easy to set up&lt;/strong&gt;: While it is not completely hands-off (you still manage the JuiceFS client, the local SQLite instance, and Litestream's replication configuration like Celesto AI does), this stack is drastically simpler to bootstrap than a distributed database cluster.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Per-agent isolation&lt;/strong&gt;: Each sandbox gets its own dedicated SQLite metadata file. This eliminates noisy-neighbor problems. Metadata contention simply does not happen across agents, because each one has its own independent "brain" living right next to its mount.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Performance that fits the scope&lt;/strong&gt;: A single agent workload typically produces tens of thousands to a few million files. SQLite handles this range comfortably with low latency and minimal resource footprint. While this approach obviously does not scale to the hundreds of billions of files that &lt;a href="https://juicefs.com/docs/cloud/" rel="noopener noreferrer"&gt;JuiceFS Enterprise Edition&lt;/a&gt; can manage globally, it falls squarely into the comfortable zone for individual agent workspaces.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Of course, this approach comes with important trade-offs and considerations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;No cross-agent sharing&lt;/strong&gt;: The file system is inherently tied to a single SQLite instance. Agents cannot mount and share the same workspace concurrently. If your use case requires collaborative access or shared datasets across sandboxes, you would need a different strategy.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Resume latency depends on object store speed&lt;/strong&gt;: When a sandbox resumes, the system must download the latest SQLite snapshot from the object store and replay the WAL frames. This adds noticeable startup time compared to a hot-mounted local disk. Tuning Litestream's &lt;code&gt;sync-interval&lt;/code&gt; involves a direct trade-off between durability guarantees and resume performance.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Single-writer limitation&lt;/strong&gt;: SQLite is fundamentally a single-writer database. Heavy parallel writes from multiple processes within the same sandbox can become a bottleneck. For autonomous agents following a largely sequential workflow, this limitation is rarely a practical issue. But it is worth keeping in mind for highly concurrent, multi-threaded agent designs.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The stack is not a universal solution. It is a deliberate, opinionated choice that trades global scalability and shared access for extreme simplicity and tight per-agent isolation. And for the AI sandbox use case, that trade-off often makes perfect sense.&lt;/p&gt;

&lt;h2&gt;
  
  
  The performance reality
&lt;/h2&gt;

&lt;p&gt;This architecture is beautiful, but it isn't magic. You are trading local SSD latency for object storage durability and infinite scale. The benchmarks reveal a clear pattern:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Small files need attention: Writing 5,000 tiny 4 KiB files to an object-backed file system involves heavy metadata churn. Even with SQLite's fast local access, every &lt;code&gt;close()&lt;/code&gt; and &lt;code&gt;fsync()&lt;/code&gt; triggers WAL writes and Litestream replication. In these scenarios, the local root disk (ext4/XFS) will outperform the durable mount by a factor of 2x to 5x. See Celesto's &lt;a href="https://celesto.ai/blog/posts/platform/petabyte-scale-storage/#many-small-files" rel="noopener noreferrer"&gt;blog post&lt;/a&gt; for more details.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Large sequential workloads shine: For datasets, model checkpoints, build artifacts, and logs (files &amp;gt; 10 MB), the object-backed mount achieves upwards of 200 MiB/s, which is nearly 90% of the performance of the local disk. The background block uploads are streamed efficiently. See Celesto's &lt;a href="https://celesto.ai/blog/posts/platform/petabyte-scale-storage/#large-files" rel="noopener noreferrer"&gt;blog post&lt;/a&gt; for more details.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The golden rule for agents is to use the object-backed mount drive for durable, long-term state reports and large artifacts. Use the local ephemeral disk for high-churn, temporary scratch work that doesn't need to survive a restart. This hybrid approach gives you infinite capacity without sacrificing performance where it truly hurts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final thoughts
&lt;/h2&gt;

&lt;p&gt;Cloudflare CFO Thomas Seifert recently predicted that within five years, non-human traffic could be as much as 1,000 times human traffic, making humans &lt;a href="https://www.theregister.com/networks/2026/08/07/humans-will-be-a-rounding-error-on-the-internet-says-cloudflare-exec/5284429" rel="noopener noreferrer"&gt;"a rounding error on the internet."&lt;/a&gt; He caveated that he's "called it wrong at every point along the way," but the direction is hard to dispute: AI and machine traffic already overtook human traffic in May 2026.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbu936hhw91rlrbfmkwuq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbu936hhw91rlrbfmkwuq.png" alt=" " width="800" height="341"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If that is true for network packets, it is likely true for storage as well.&lt;/p&gt;

&lt;p&gt;Every autonomous agent session produces a workspace. Multiply that by millions of agents running continuously, resuming, retrying, and branching. Human developers generate storage in bursts tied to working hours. Agents generate it constantly, at machine speed.&lt;/p&gt;

&lt;p&gt;This raises an uncomfortable question: if non-human consumers are about to dominate storage demand, are we designing for the right consumer? The combo of JuiceFS + SQLite + Litestream is one answer, but not the only one. Other approaches worth watching include object storage as the primary interface, shared metadata engine with concurrent access at agent scale (i.e., the JuiceFS Enterprise Edition), copy-on-write workspace snapshots, and &lt;a href="https://juicefs.com/en/blog/solutions/ai-agent-compare-agentfs-tigerfs-juicefs" rel="noopener noreferrer"&gt;many others&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Portable, shareable, scalable file storage for AI agents is still in its opening moves.&lt;/p&gt;

&lt;p&gt;If you have any feedback on this article or ideas to share, we invite you to participate in the &lt;a href="https://github.com/juicedata/juicefs/discussions/" rel="noopener noreferrer"&gt;discussions on GitHub&lt;/a&gt; and join &lt;a href="http://go.juicefs.com/discord" rel="noopener noreferrer"&gt;our community on Discord&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
    <item>
      <title>Keeping GPUs Fed: How Meta's AI Storage Architecture Aligns with JuiceFS</title>
      <dc:creator>DASWU</dc:creator>
      <pubDate>Fri, 18 Sep 2026 07:18:00 +0000</pubDate>
      <link>https://dev.to/daswu/keeping-gpus-fed-how-metas-ai-storage-architecture-aligns-with-juicefs-55d1</link>
      <guid>https://dev.to/daswu/keeping-gpus-fed-how-metas-ai-storage-architecture-aligns-with-juicefs-55d1</guid>
      <description>&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;Recently, I came across an eye-opening &lt;a href="https://engineering.fb.com/2026/07/01/data-infrastructure/metas-ai-storage-blueprint-at-scale/" rel="noopener noreferrer"&gt;blog post&lt;/a&gt; and &lt;a href="https://www.youtube.com/watch?v=xc_BxLsijWA" rel="noopener noreferrer"&gt;presentation&lt;/a&gt; from Meta's engineering team detailing their AI storage blueprint at scale. Meta operates hundreds of exabyte-scale storage clusters serving everything from Facebook and Instagram to Meta AI.&lt;/p&gt;

&lt;p&gt;What caught my attention wasn't just the scale—it was the architecture. After years of wrestling with AI's brutal latency and throughput demands, Meta has converged on a clear pattern: &lt;strong&gt;data-metadata separation, rich client design, and comprehensive caching.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here's the fascinating part: this blueprint isn't unique to Meta. &lt;a href="https://juicefs.com/docs/community/architecture" rel="noopener noreferrer"&gt;&lt;strong&gt;JuiceFS&lt;/strong&gt;&lt;/a&gt; &lt;strong&gt;shares these same first principles of design.&lt;/strong&gt; The hyperscaler's new construction closely resembles the cloud-native architecture that JuiceFS has advocated for years.&lt;/p&gt;

&lt;p&gt;In this post, I'll analyze both architectures, compare their core components, dive into their cross-region capabilities, and explain why this convergence in architecture makes the most sense for a distributed file system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why storage matters for AI workloads
&lt;/h2&gt;

&lt;p&gt;Before diving into the architectures, it's worth understanding why storage has become such a critical component for AI infrastructure.&lt;/p&gt;

&lt;p&gt;Meta puts it succinctly: &lt;em&gt;"If AI is the brain, storage is the memory: capability and speed are highly dependent on the size of memory and speed of retrieval."&lt;/em&gt; Yet while AI compute performance has roughly tripled every two years, storage and interconnect performance growth have been far more modest. This disparity has made storage bottlenecks one of the primary contributors to GPU stalls for AI workloads, which is directly impacting both expenditures and time to market.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0cy9ogyjau7wbr2dag5p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0cy9ogyjau7wbr2dag5p.png" alt=" " width="800" height="432"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The math is brutal. During model training, hundreds of thousands of GPUs iterate over vast datasets across multiple epochs. GPUs synchronize their state periodically. If a single GPU stalls due to slow storage, it slows down the entire training job. In Meta's presentation, if GPUs were idle for about 20% of the time, &lt;em&gt;"that 20% stall means about tens of millions of dollars loss every hour"&lt;/em&gt; for a data center the size of Manhattan.&lt;/p&gt;

&lt;p&gt;Beyond the direct dollar cost, there is an even more consequential dimension: opportunity cost. The AI industry is moving at an unprecedented pace. Major model releases, which took roughly 104 weeks between 2020 and 2022, have now been compressed to approximately 4-week cycles by 2026. Every week of delay in shipping a frontier model translates to lost market share, diminished competitive advantage, and missed revenue opportunities.&lt;/p&gt;

&lt;p&gt;In short, storage is what makes the difference between a model shipping in weeks versus months and between a billion-dollar data center running at 80% versus 100% efficiency.&lt;/p&gt;

&lt;h2&gt;
  
  
  Meta's AI storage blueprint: the evolved architecture
&lt;/h2&gt;

&lt;p&gt;After iterating on their storage infrastructure, Meta has arrived at a definitive architectural blueprint to solve the GPU stall problem at its root while continuing to support existing external and internal products. The training stack has been gradually migrating to the BLOB-storage interface. The motivation is more about performance: AI workloads demand predictable, bounded pMax latency. Traditional layered metadata lookups could add hundreds of milliseconds per request, stalling GPUs during training. By moving to a BLOB-centric interface, Meta leverages flash drives to deliver the low, predictable latencies required to keep thousands of GPUs productive.&lt;/p&gt;

&lt;h3&gt;
  
  
  The three-component architecture
&lt;/h3&gt;

&lt;p&gt;The diagram below illustrates the request flow for Meta's BLOB storage &lt;code&gt;getObject&lt;/code&gt; API and shows the system's three major components.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp1w8cm6dl32e6fwns8xi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fp1w8cm6dl32e6fwns8xi.png" alt=" " width="800" height="609"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The Tectonic data storage&lt;/strong&gt;: At the foundation sits Tectonic, a regional, multi-tenant storage fabric providing high durability and availability through erasure coding. It supports tiering across HDDs and flash drives, intelligently managing hot, cold, and warm data placement.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The BLOB server and metadata layer&lt;/strong&gt;: Operating on top of Tectonic, the BLOB layer handles API requests, stores metadata, and exposes a global, infinitely scalable storage fabric with policies allowing users to trade off between durability and availability, abstracting away the complexity of the underlying block storage.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The rich client SDK&lt;/strong&gt;: Meta eliminated the data-plane proxy and built a rich client SDK with a Tectonic BlockClient embedded within it, enabling direct data reading from Tectonic storage servers to clients. The client also implements hedged reads and dynamic concurrency control to reduce tail latencies.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To maximize cache efficiency, Meta also leverages spare memory on GPU hosts as a distributed data cache by integrating peers from the Owl subsystem directly into the BLOB-storage client SDK, which yields an average cache hit rate of 80%.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cross-region data access with tiered caching
&lt;/h3&gt;

&lt;p&gt;Meta's cross-region approach treats storage as a global data lake, but with tiered caches accelerating access across geographical boundaries.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fls2lesw0wzyi4eyh3c5o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fls2lesw0wzyi4eyh3c5o.png" alt=" " width="800" height="536"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cache Tier&lt;/th&gt;
&lt;th&gt;Location&lt;/th&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;L1 Cache&lt;/td&gt;
&lt;td&gt;GPU host memory (RAM)&lt;/td&gt;
&lt;td&gt;Hottest, most frequently accessed data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L2 Cache&lt;/td&gt;
&lt;td&gt;GPU host flash (SSD)&lt;/td&gt;
&lt;td&gt;Local flash tier for warm data&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;L3 Cache&lt;/td&gt;
&lt;td&gt;Regional flash-based BLOB storage&lt;/td&gt;
&lt;td&gt;Regional cache before falling back to global BLOB storage&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The results speak for themselves: average ingestion time dropped by 93%, and worst-case ingestion time dropped by 97%. This is what it looks like when storage stops being the bottleneck and starts enabling AI research velocity.&lt;/p&gt;

&lt;h2&gt;
  
  
  JuiceFS: cloud-native, AI-native from day one
&lt;/h2&gt;

&lt;p&gt;For readers already familiar with JuiceFS, the parallels to Meta's design are immediately apparent. The architecture Meta converged upon after years of iteration shares the same core principles that have guided JuiceFS from day one.&lt;/p&gt;

&lt;p&gt;JuiceFS was architected from the ground up on the data-metadata separation principle, which also consists of three core components:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Object storage&lt;/strong&gt;: JuiceFS uses object storage as the reliable, scalable backend for actual data storage. It supports all major public cloud offerings (AWS S3, Alibaba OSS, Google Cloud Storage, etc.) as well as self-hosted solutions like Ceph and MinIO. This layer ensures that AI datasets, which can be as large as petabytes or exabytes, are safe, accessible, and can grow almost without limit.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Metadata engine&lt;/strong&gt;: The metadata engine stores all file system metadata, including file names, sizes, permissions, directory structures, and the mapping between files and their underlying data blocks. JuiceFS supports multiple engine options, including Redis, MySQL, PostgreSQL, TiKV, and many others. For extremely large-scale deployments, &lt;a href="https://juicefs.com/docs/cloud/" rel="noopener noreferrer"&gt;JuiceFS Enterprise Edition&lt;/a&gt; offers a proprietary Raft-based distributed metadata engine that provides high availability and strong consistency across multiple nodes.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;JuiceFS client&lt;/strong&gt;: The client handles all file I/O logic, including data slicing, merging, caching, and POSIX semantics enforcement. It communicates with both the metadata engine and object storage directly, coordinating reads and writes intelligently.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Cross-region strategy: mirrors and replication
&lt;/h3&gt;

&lt;p&gt;The JuiceFS Enterprise Edition also provides a native approach to support cross-region and cross-cloud data access:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://juicefs.com/docs/cloud/guide/mirror/" rel="noopener noreferrer"&gt;&lt;strong&gt;Mirror file systems&lt;/strong&gt;&lt;/a&gt;: Deploy a full metadata cluster and optionally an object storage replica in a mirror region. Metadata syncs automatically from the source region with approximately 1-second latency under normal network conditions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://juicefs.com/docs/cloud/guide/replication/" rel="noopener noreferrer"&gt;&lt;strong&gt;Async data replication&lt;/strong&gt;&lt;/a&gt;: JuiceFS supports cross-region and cross-cloud asynchronous replication with a one-to-many topology. Data written in the primary region is replicated to multiple mirror regions in the background, enabling global datasets without manual copying.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://juicefs.com/docs/cloud/guide/distributed-cache/" rel="noopener noreferrer"&gt;&lt;strong&gt;Distributed cache&lt;/strong&gt;&lt;/a&gt;: JuiceFS clients on mounting hosts can form a distributed cache group using a consistent hashing ring, sharing cached data blocks with one another. This is particularly effective for model training, where datasets are repeatedly accessed across GPU nodes.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Smart read strategy&lt;/strong&gt;: Clients preferentially read from the local region. If the data isn't fully synced yet, they seamlessly fall back to the source region.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fykw58o3gxjimby7r65is.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fykw58o3gxjimby7r65is.png" alt=" " width="481" height="402"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This design resonates with Meta's global data lake + tiered cache approach but offers a more flexible, cloud-native model that works across public clouds, private data centers, and hybrid environments.&lt;/p&gt;

&lt;h2&gt;
  
  
  Architectural similarities
&lt;/h2&gt;

&lt;p&gt;On the surface, Meta mainly exposes a BLOB-style interface, while JuiceFS primarily provides a POSIX file system. Yet beneath these protocol differences, the architectural parallels are noteworthy. Both systems converged on the same three-component design with comprehensive caching and cross-region capabilities:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;Meta's BLOB-Storage&lt;/th&gt;
&lt;th&gt;JuiceFS&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Data storage layer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Tectonic: regional block storage fabric with erasure coding, tiering across HDDs and flash&lt;/td&gt;
&lt;td&gt;Object storage: S3, OSS, GCS, Ceph, MinIO, and others&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Metadata layer&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Unified metadata scheme backed by ZippyDB with O(1) chunk lookups&lt;/td&gt;
&lt;td&gt;Metadata engine: Redis, MySQL, PostgreSQL, TiKV, JuiceFS Enterprise Edition&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Rich client&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;"Fat client SDK" with embedded Tectonic BlockClient, streaming directly from storage servers&lt;/td&gt;
&lt;td&gt;JuiceFS client: handles all I/O logic, slicing, merging, caching; supports FUSE, POSIX, Hadoop SDK, CSI, S3 Gateway&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Distributed cache&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Tiered caching layers&lt;/td&gt;
&lt;td&gt;Memory + local cache + distributed cache group&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cross-region&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Tiered caching with global BLOB storage data lakes&lt;/td&gt;
&lt;td&gt;Primary region + mirror file systems&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  A concrete example: explicit cache hydration
&lt;/h3&gt;

&lt;p&gt;One interesting example is how both systems handle proactive cache hydration, which is the practice of loading data into the cache before it's actually needed.&lt;/p&gt;

&lt;p&gt;Meta exposes a &lt;code&gt;prefetch()&lt;/code&gt; API through the BLOB-storage SDK, while JuiceFS provides a &lt;a href="https://juicefs.com/docs/cloud/reference/command_reference/#warmup" rel="noopener noreferrer"&gt;&lt;code&gt;warmup&lt;/code&gt; &lt;/a&gt; subcommand. Both serve the same purpose: allowing the client side to signal to the storage system exactly which data will be needed in the near future so it can be preloaded into cache before the GPU finishes its current work.&lt;/p&gt;

&lt;p&gt;This is especially critical in distributed training, where thousands of GPUs may simultaneously read the same dataset. Without cache hydration, the initial requests would have to read from the underlying storage (Tectonic for BLOB and object store for JuiceFS) with higher latency. With &lt;code&gt;prefetch()&lt;/code&gt; or &lt;code&gt;juicefs warmup&lt;/code&gt;, data is staged in the distributed cache before the GPUs need it. In essence, both systems recognize the limitations of network-based storage and passive caching, and they both provide explicit prefetching as an important tool to compensate for what the storage layer alone cannot guarantee.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Meta's AI storage blueprint is a testament to the architectural patterns that emerge when a hyperscaler evolves its data infrastructure to hundreds of exabytes. After years of iteration, they arrived at a design that's similar to what JuiceFS has championed from day one: data-metadata separation, a rich client that handles I/O logic directly, tiered caching, and cross-region capabilities.&lt;/p&gt;

&lt;p&gt;The same architectural pattern is also available to teams of any size. JuiceFS brought it into the big data era and now brings it into the AI age, with the flexibility of multi-cloud cross-region strategies, distributed caching, and a rich ecosystem of access protocols.&lt;/p&gt;

&lt;p&gt;Hopefully, storage is no longer the bottleneck in your AI training stack. At least, it doesn't have to be.&lt;/p&gt;

&lt;p&gt;If you have any feedback on this article or ideas to share, we invite you to participate in the &lt;a href="https://github.com/juicedata/juicefs/discussions/" rel="noopener noreferrer"&gt;discussions on GitHub&lt;/a&gt; and join &lt;a href="http://go.juicefs.com/discord" rel="noopener noreferrer"&gt;our community on Discord&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Giving AI Agents a Real File System</title>
      <dc:creator>DASWU</dc:creator>
      <pubDate>Fri, 21 Aug 2026 08:03:03 +0000</pubDate>
      <link>https://dev.to/daswu/giving-ai-agents-a-real-file-system-2a2j</link>
      <guid>https://dev.to/daswu/giving-ai-agents-a-real-file-system-2a2j</guid>
      <description>&lt;h2&gt;
  
  
  The new root user
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://en.wikipedia.org/wiki/Everything_is_a_file" rel="noopener noreferrer"&gt;&lt;em&gt;"Everything is a file."&lt;/em&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For decades, this single sentence defined the core Unix philosophy. Today, the primary operator interacting with those files is no longer a human at a terminal but an autonomous LLM agent spinning up environments, editing code, and executing scripts.&lt;/p&gt;

&lt;p&gt;Because LLMs were pre-trained on codebases, wiki pages, documentation, directory trees, and terminal logs, POSIX file system operations serve as the zero-friction API for agentic reasoning. Giving an agent access to a familiar workspace lets it leverage standard developer tools immediately.&lt;/p&gt;

&lt;p&gt;The following is an example output of an AI agent performing code reviews:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Verified, the following command returns only the method definition.&lt;/span&gt;
&lt;span class="c"&gt;# Your change doesn't have any actual invocation of this method.&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="s2"&gt;"store.addDelayedStaging"&lt;/span&gt; &lt;span class="nb"&gt;.&lt;/span&gt; &lt;span class="nt"&gt;--include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"*.go"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;However, while a classic POSIX interface is necessary, modern AI agents demand far more from their storage than standard local file systems were built to deliver. To support a massive number of ephemeral, fast-moving agent sessions, the underlying storage layer must evolve as well.&lt;/p&gt;

&lt;h3&gt;
  
  
  Architectural paradox
&lt;/h3&gt;

&lt;p&gt;This shift introduces a fundamental storage paradox for AI infrastructure engineers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;At the surface layer&lt;/strong&gt;: The agent requires a workspace that strictly behaves like a classic, local POSIX file system so that standard developer tools (&lt;code&gt;git&lt;/code&gt;, &lt;code&gt;python&lt;/code&gt;, &lt;code&gt;grep&lt;/code&gt;, compilers, etc.) function seamlessly without custom glue code.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Under the hood&lt;/strong&gt;: Traditional instance-bound storage breaks down when managing hundreds or thousands of ephemeral, short-lived, or parallel agent sessions. The underlying backend must provide remote snapshotting, consistent multi-agent concurrent access, elasticity, and scalability while still being compatible with that familiar file interface.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To build scalable infrastructure for autonomous AI agents, we must preserve the classic Unix interface while completely re-engineering the storage engine beneath it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why AI agents love the file system
&lt;/h2&gt;

&lt;p&gt;If APIs are rigid software contracts, POSIX file systems are the open playground. Giving an LLM agent direct access to a file system workspace unlocks operational advantages that structured tool calls simply cannot match.&lt;/p&gt;

&lt;h3&gt;
  
  
  Offloading the context window
&lt;/h3&gt;

&lt;p&gt;Even the most advanced models with million-token context windows experience severe latency, high costs, and &lt;a href="https://www.trychroma.com/research/context-rot" rel="noopener noreferrer"&gt;performance or attention degradation&lt;/a&gt; when a massive repository or gigabytes of build logs are included in an LLM prompt. The file system serves as secondary memory, allowing the agent to store raw state externally and selectively read, seek, or stream only the relevant fragments it needs at runtime.&lt;/p&gt;

&lt;h3&gt;
  
  
  The universal abstraction
&lt;/h3&gt;

&lt;p&gt;A file system requires zero protocol overhead when an agent has shell execution access, unlike external services or database integrations that need custom JSON schemas or Model Context Protocol (MCP) tool wrappers. LLMs already possess a profound understanding of standard file operations and directory navigation out of the box. Exposing a POSIX file system gives the agent a universal workspace across any language, framework, or file format, without needing to define or maintain a single custom tool schema.&lt;/p&gt;

&lt;h3&gt;
  
  
  Unlocking the Unix ecosystem
&lt;/h3&gt;

&lt;p&gt;By granting an agent a file system, it immediately inherits decades of battle-tested CLI tools: &lt;code&gt;grep&lt;/code&gt;, &lt;code&gt;awk&lt;/code&gt;, &lt;code&gt;git&lt;/code&gt;, &lt;code&gt;sed&lt;/code&gt;, compilers, and static analyzers. Rather than inventing "AI-native code search," an agent can simply execute grep or analyze git diff using tools already optimized for speed and scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  Database-driven metadata meets elastic storage
&lt;/h2&gt;

&lt;p&gt;When building agentic platforms, the naive approach is to rely on local block storage attached directly to a compute instance (e.g., standard ephemeral instance disks or container volumes). While simple initially, instance-bound disks rapidly fail under the demands of autonomous agents:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Ephemeral state loss&lt;/strong&gt;: When an agent container scales down, crashes, or migrates to another host, all workspace state and context stored on instance-tied disks are lost unless manually synced.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Concurrent data access&lt;/strong&gt;: Multi-agent workflows often require sub-agents (e.g., a planner, coder, and reviewer) to operate on the same repository simultaneously. Instance-attached disks cannot safely allow multi-node concurrent reads and writes without file lock corruption or race conditions.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Non-native data backup&lt;/strong&gt;: Compute-instance disks lack built-in data backup or snapshotting, which makes rolling back the environment after an agent error slow and expensive.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To overcome these constraints, modern agent file architectures decouple the &lt;strong&gt;metadata engine&lt;/strong&gt; (directory trees, permissions, file locks, and state) from the &lt;strong&gt;data storage layer&lt;/strong&gt; (file payloads).&lt;/p&gt;

&lt;p&gt;By backing the metadata layer with specialized database engines and offloading raw file blocks to elastic object or remote storage, file system operations essentially become database queries. The agent gets the exact POSIX semantics its tools expect, while the infrastructure gains consistency guarantees, automatic backups, and infinite horizontal scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ecosystem landscape
&lt;/h2&gt;

&lt;p&gt;Innovative projects across the storage industry demonstrate how database backends are powering this new era of AI agent file storage. Here, we are exploring a few notable approaches.&lt;/p&gt;

&lt;h3&gt;
  
  
  Lightweight embedded approach: AgentFS with SQLite
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://github.com/tursodatabase/agentfs" rel="noopener noreferrer"&gt;AgentFS&lt;/a&gt;, developed by Turso, embodies the lightweight embedded approach within the database-as-file-system paradigm. Designed specifically for agent sandboxing, it stores every piece of runtime state (file operations, key-value entries, and tool call histories) in a single SQLite database. The underlying engine is Turso, a full rewrite of SQLite in Rust that adds native asynchronous I/O and concurrent access, with optional cloud sync for portability across environments.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd6hzgxtgn7a2mnd7w6j5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd6hzgxtgn7a2mnd7w6j5.png" alt=" " width="800" height="605"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;AgentFS exposes a POSIX-like virtual file system via an SDK or a FUSE mount. This unified storage foundation delivers copy-on-write isolation for safe agent execution, instant SQL-queryable audit trails for debugging and compliance, and single-file snapshots that make agent state trivial to share, version, and reproduce.&lt;/p&gt;

&lt;h3&gt;
  
  
  Client-server transactional approach: TigerFS with PostgreSQL
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://github.com/timescale/tigerfs" rel="noopener noreferrer"&gt;TigerFS&lt;/a&gt;, developed by Timescale, mounts a PostgreSQL database as a transactional file system over FUSE (Linux) or NFS (macOS). Every file maps to a row, directories map to tables, and file contents become columns. Every write is a full ACID transaction, enabling multiple agents and humans to read and write the same files concurrently with true ACID guarantees.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdv09s8s175sh04blpbkf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdv09s8s175sh04blpbkf.png" alt=" " width="800" height="136"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Multi-agent task coordination becomes as simple as moving a file (&lt;code&gt;`mv task.md ./doing/`&lt;/code&gt;), with PostgreSQL transactions natively preventing race conditions without custom API orchestration. TigerFS offers two modes: file‑first for workspaces with atomic writes and reversible savepoints, and data‑first for exploring existing databases with Unix tools like &lt;code&gt;`ls`&lt;/code&gt;, &lt;code&gt;`cat`&lt;/code&gt;, and &lt;code&gt;`grep`&lt;/code&gt;, with filters pushed down as SQL. Every change is versioned and fully reversible, and the file system itself serves as the API.&lt;/p&gt;

&lt;h3&gt;
  
  
  Decoupled distributed approach: JuiceFS with object storage
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://github.com/juicedata/juicefs" rel="noopener noreferrer"&gt;JuiceFS&lt;/a&gt; takes a fundamentally different architectural path: it decouples metadata and data storage entirely. File data is split into chunks and stored in S3-compatible object storage, while metadata lives in a separate, pluggable engine, supporting Redis, MySQL, PostgreSQL, TiKV, SQLite, JuiceFS Enterprise Edition (EE) metadata engine, and more. This separation enables JuiceFS to deliver strict POSIX compliance alongside shared, multi-node concurrent access, allowing hundreds of agent instances to safely work across a shared workspace. JuiceFS can be connected using FUSE, Kubernetes CSI, or SDKs (like Python, Hadoop, and S3 gateway), which makes it a flexible storage option for agents as well as other AI processes, including data preparation, model training, and deployment in different regions. Every committed change is immediately visible across all mounts with strong consistency, and the system scales elastically with object storage economics.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffi3nsi2gdqootw5b12nq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffi3nsi2gdqootw5b12nq.png" alt=" " width="800" height="605"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Together, projects like AgentFS, TigerFS, and JuiceFS represent distinct architectural explorations in the quest to provide agents with a persistent, shareable file system. Each tackles the problem from a different angle, yet all share a common recognition: agents naturally work with files. There is no clear consensus on which approach will become the dominant standard, and perhaps no single winner exists. What is certain is that the file system is re-emerging as a core primitive for AI agent orchestration, precisely because the POSIX interface aligns so naturally with how agents read, write, and organize data, and ultimately because agents complete tasks by acting on data and files.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agents' full toolbox
&lt;/h2&gt;

&lt;p&gt;While a database-backed file system provides a zero-friction home for code and context, a file system alone isn't enough for complex enterprise workflows. Real-world agents also need to orchestrate across multiple systems of record, for example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The file system&lt;/strong&gt;: As discussed above, the file system serves as the primary workspace for source code, local logs, build artifacts, and raw data manipulation.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Databases &amp;amp; APIs&lt;/strong&gt;: The structured state layer for transaction logs, vector search embeddings, and relational data that models need to query or update.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The Model Context Protocol (MCP)&lt;/strong&gt;: An open standard protocol that standardizes how LLM applications connect to external data sources and tools through MCP servers. It does not replace existing APIs. Instead, it provides a common interface and transport layer on top of them so clients can discover and use capabilities without custom integrations for every application or model.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;By combining elastic file workspaces with standardized protocols like MCP, developers can build agents that are autonomous, practical, and powerful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final thoughts
&lt;/h2&gt;

&lt;p&gt;We aren't just building sandboxes anymore. Instead, we are designing a new type of operating system. In this model, the LLM acts as the CPU, the context window is the fast L1 cache, and database-backed file systems serve as persistent NVMe drives.&lt;/p&gt;

&lt;p&gt;The industry is still actively debating what an agent's ultimate memory architecture should look like: whether it should rely on vector databases, structured state stores, or file-based paradigms. While the ideal abstraction may continue to evolve, the file system is already proving to be a core foundation in agent infrastructure. Let's honor the classic Unix philosophy and give agents seamless access to decades of battle-tested software tools through a familiar POSIX interface and the performance and scale that modern distributed backends deliver.&lt;/p&gt;

&lt;p&gt;If you have any feedback on this article or ideas to share, we invite you to participate in the &lt;a href="https://github.com/juicedata/juicefs/discussions/" rel="noopener noreferrer"&gt;discussions on GitHub&lt;/a&gt; and join &lt;a href="http://go.juicefs.com/discord" rel="noopener noreferrer"&gt;our community on Discord&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>GPFS vs. Alluxio vs. JuiceFS: Architecture and Use Cases Compared</title>
      <dc:creator>DASWU</dc:creator>
      <pubDate>Fri, 14 Aug 2026 03:58:21 +0000</pubDate>
      <link>https://dev.to/daswu/gpfs-vs-alluxio-vs-juicefs-architecture-and-use-cases-compared-4bm6</link>
      <guid>https://dev.to/daswu/gpfs-vs-alluxio-vs-juicefs-architecture-and-use-cases-compared-4bm6</guid>
      <description>&lt;p&gt;In AI workloads, scenarios such as training, inference, model distribution, agents, and data lakes each have distinct requirements for throughput, latency, concurrent access, POSIX compatibility, consistency, cost, and operational complexity.&lt;/p&gt;

&lt;p&gt;In this article, we’ll first examine the typical I/O patterns of AI workloads, identifying the practical challenges these workloads pose to storage systems. Then, we’ll compare General Parallel File System (GPFS), Alluxio, and &lt;a href="https://juicefs.com/docs/community/introduction/" rel="noopener noreferrer"&gt;JuiceFS&lt;/a&gt;, analyzing their architectural differences, file system semantics, caching mechanisms, cost models, and applicability across different AI scenarios.&lt;/p&gt;

&lt;h2&gt;
  
  
  I/O patterns and storage challenges in AI workloads
&lt;/h2&gt;

&lt;p&gt;Based on our interactions with enterprises across various AI domains, we can categorize common requirements as follows.&lt;/p&gt;

&lt;h3&gt;
  
  
  Autonomous driving: large‑scale data generation and training
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://en.wikipedia.org/wiki/Self-driving_car" rel="noopener noreferrer"&gt;Autonomous driving&lt;/a&gt; is one of the most data‑intensive AI storage scenarios. Large fleets of data‑collection vehicles continuously generate images, videos, sensor data, and trajectory logs, which undergo cleaning, labeling, and format conversion before entering model training pipelines. Common data formats include .mcap, .pack, TFRecord, and LMDB. These environments require highly stable and scalable storage due to the long data pipelines, diverse formats, and heavy training workloads.&lt;/p&gt;

&lt;h3&gt;
  
  
  LLM training: full‑pipeline data access
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://en.wikipedia.org/wiki/Foundation_model" rel="noopener noreferrer"&gt;Foundation model&lt;/a&gt; scenarios cover multiple stages, including data cleaning, model training, checkpoint reads/writes, and inference services. Model weights, training data, and intermediate results are accessed repeatedly across different stages. The storage system must sustain long‑running jobs while maintaining stable data access during task failures, node outages, and training recovery.&lt;/p&gt;

&lt;h3&gt;
  
  
  Multimodal models: coexistence of small files, aggregated data, and model files
&lt;/h3&gt;

&lt;p&gt;AIGC scenarios include text‑to‑image, image‑to‑image, text‑to‑video, image‑to‑video, and 3D generation. Training inputs may consist of many individual images or video clips, or they may be aggregated into formats like LMDB or Parquet to improve training efficiency. The training process also generates checkpoints and outputs model files such as safetensors. This makes the data landscape more complex than in single‑task training scenarios.&lt;/p&gt;

&lt;h3&gt;
  
  
  Compute platforms: multi‑cloud collaboration and model distribution
&lt;/h3&gt;

&lt;p&gt;Compute platforms focus on distributing models and data across multiple environments. Users may pull models from external repositories or upload their own, and then run training or inference workloads across different clusters and cloud environments. The key challenge is reducing redundant copying while enabling consistent access to the same data across different compute environments.&lt;/p&gt;

&lt;h3&gt;
  
  
  Quantitative finance: balancing performance and cost
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://en.wikipedia.org/wiki/Quantitative_analysis_(finance)" rel="noopener noreferrer"&gt;Quantitative finance&lt;/a&gt; typically involves smaller data volumes than autonomous driving or AIGC. However, as techniques such as Transformer‑based models, neural network training, time‑series modeling, and market graph structure analysis are increasingly adopted, storage cost is becoming a more explicit selection factor.  &lt;/p&gt;

&lt;p&gt;In recent discussions with quantitative finance customers, we’ve observed growing attention to storage costs. High‑performance parallel file systems like GPFS are inherently performance‑oriented, especially in large‑capacity, all‑flash configurations. In some on‑premises deployments, the storage investment for a large‑capacity all‑flash GPFS cluster can approach the cost of an entire 5090 GPU cluster.&lt;/p&gt;

&lt;h3&gt;
  
  
  AI agents: data sharing in short‑lived sandboxes
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.ibm.com/think/topics/ai-agents" rel="noopener noreferrer"&gt;AI agents&lt;/a&gt; represent a rapidly emerging scenario. They often involve many short‑lived sandboxes—each executing a subtask with a lifetime of just seconds or even less.  &lt;/p&gt;

&lt;p&gt;While these tasks run briefly, the context, model files, tool files, and intermediate results must be shared among subtasks. If each sandbox independently mounts a file system, and the mount process itself takes several seconds, task scheduling efficiency suffers. A more practical approach is to pre‑mount the file system on the host and then expose it to sandboxes via bind mounts, PVCs, or similar mechanisms.  &lt;/p&gt;

&lt;p&gt;From a storage perspective, AI agents are less concerned with sheer capacity and more focused on data continuity, shared access, and mount efficiency within short‑lived tasks. As agent complexity grows, demands on file system semantics and data sharing capabilities will continue to rise.  &lt;/p&gt;

&lt;p&gt;The table below shows typical I/O patterns and storage requirements for AI workloads:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Typical I/O pattern&lt;/th&gt;
&lt;th&gt;Key storage challenges&lt;/th&gt;
&lt;th&gt;Selection criteria&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Autonomous driving&lt;/td&gt;
&lt;td&gt;Large‑file throughput, mmap random reads, small‑file reads&lt;/td&gt;
&lt;td&gt;Large data scale, long training pipelines, high random‑read pressure&lt;/td&gt;
&lt;td&gt;Throughput, caching, metadata capabilities, capacity cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LLM training&lt;/td&gt;
&lt;td&gt;Large‑file reads/writes, mixed reads, checkpoint reads/writes&lt;/td&gt;
&lt;td&gt;Full‑pipeline access, high stability requirements for long‑running tasks&lt;/td&gt;
&lt;td&gt;Stable throughput, concurrent access, fault recovery&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multimodal models&lt;/td&gt;
&lt;td&gt;Small‑file reads, aggregated large‑file reads, model file access&lt;/td&gt;
&lt;td&gt;Coexistence of small and large files, significant multi‑task concurrency&lt;/td&gt;
&lt;td&gt;Caching, metadata management, multi‑task concurrency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compute platforms&lt;/td&gt;
&lt;td&gt;Model distribution, cross‑cluster access, multi‑cloud collaboration&lt;/td&gt;
&lt;td&gt;Data must be consistently accessible across environments&lt;/td&gt;
&lt;td&gt;Unified namespace, multi‑cloud distribution, cache governance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quantitative finance&lt;/td&gt;
&lt;td&gt;Large‑file sequential reads, small‑file reads, training/backtesting access&lt;/td&gt;
&lt;td&gt;Rising cost sensitivity, need to balance performance and cost&lt;/td&gt;
&lt;td&gt;Capacity cost, scalability, long‑term operations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AI agents&lt;/td&gt;
&lt;td&gt;Small I/O, multi‑client sharing, short‑lived access&lt;/td&gt;
&lt;td&gt;Mount efficiency, data continuity, task isolation&lt;/td&gt;
&lt;td&gt;File system semantics, shared access, mount methods&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  GPFS vs. JuiceFS
&lt;/h2&gt;

&lt;h3&gt;
  
  
  From PFS to GPFS: capabilities and limitations of parallel file systems
&lt;/h3&gt;

&lt;p&gt;To understand GPFS, it helps to first understand &lt;a href="https://www.datacore.com/glossary/parallel-file-systems/" rel="noopener noreferrer"&gt;parallel file systems&lt;/a&gt; (PFSs). Intuitively, a parallel file system separates metadata and data, enabling multiple clients to access underlying storage resources in parallel. In this architecture, metadata and data travel different paths. Clients do not need to funnel all I/O through a single node; they can concurrently access underlying disks or storage nodes. This allows hundreds of compute nodes to simultaneously read and write block devices or storage resources, breaking the network bottleneck of a single path and enabling horizontal scaling.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F20l0li15rkv33kddnpfp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F20l0li15rkv33kddnpfp.png" alt=" " width="799" height="448"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;GPFS, Lustre, and BeeGFS are all typical examples of parallel file systems.&lt;/p&gt;

&lt;p&gt;GPFS, later renamed IBM Storage Scale, is a mature parallel file system that has been widely used in high-performance computing environments for many years. It provides strong advantages in throughput, concurrent access, and consistency, making it suitable for workloads that require high performance and reliability.&lt;/p&gt;

&lt;p&gt;However, as AI infrastructure continues to scale and organizations increasingly prioritize cost optimization, the high cost and operational complexity of GPFS can become important considerations during storage selection.&lt;/p&gt;

&lt;p&gt;Typical GPFS workloads include quantitative finance, genomic analysis, physics simulations, and weather forecasting.&lt;/p&gt;

&lt;h3&gt;
  
  
  Architecture advantages and trade-offs: metanode, token locks, and strong consistency
&lt;/h3&gt;

&lt;p&gt;The architectural advantages of GPFS mainly come from two mechanisms:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Metanode&lt;/strong&gt;, which affects metadata coordination
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Distributed token locking&lt;/strong&gt;, which determines how consistency is maintained during concurrent access&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Metanode-based metadata coordination&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Under the Metanode model, a GPFS cluster usually consists of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;I/O servers
&lt;/li&gt;
&lt;li&gt;data disks
&lt;/li&gt;
&lt;li&gt;metadata disks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Metadata is stored on dedicated metadata disks. Unlike systems with centralized metadata services, GPFS allows clients to participate in parts of metadata coordination.&lt;/p&gt;

&lt;p&gt;In other words, GPFS assigns different clients as coordination nodes for different files or inodes instead of routing all metadata requests through a fixed centralized service.&lt;/p&gt;

&lt;p&gt;Therefore, GPFS clients need to maintain communication with each other, typically through port 1191.&lt;/p&gt;

&lt;p&gt;When network issues or connection failures occur, the cluster manager must determine node status and remove failed nodes from the cluster to prevent split-brain conditions and data inconsistency.&lt;/p&gt;

&lt;p&gt;This design helps distribute metadata coordination pressure. Under stable network and storage conditions, it can support high levels of concurrent access.&lt;/p&gt;

&lt;p&gt;However, it also introduces higher environmental requirements. If network quality is unstable or storage latency increases, coordination efficiency may degrade. In severe cases, users may experience:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Request stalls
&lt;/li&gt;
&lt;li&gt;Long waiters
&lt;/li&gt;
&lt;li&gt;Recovery operations requiring node restarts&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These problems do not necessarily indicate a weakness of GPFS itself. Instead, they reflect that high-performance, strongly consistent architectures require reliable networking, disk, and cluster state management.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Distributed token locks&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Distributed token locking is another key mechanism behind GPFS consistency.&lt;/p&gt;

&lt;p&gt;GPFS manages read and write operations through tokens:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Read operations require read tokens
&lt;/li&gt;
&lt;li&gt;Write operations require write tokens&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When multiple clients access the same file, GPFS controls concurrent access by granting, revoking, and transferring tokens.&lt;/p&gt;

&lt;p&gt;For example, if Client A holds a write token for a file and Client B wants to read or modify the same file, GPFS must revoke the token from A. A must flush dirty data to disk and release the token before B can proceed. This ensures strong consistency, but it also relies on stable, fast network and disk responses.&lt;/p&gt;

&lt;p&gt;If disk writes become slow or network communication is interrupted during token revocation, long waiters can appear. In practice, requests waiting on Revoke Token or Reopen Token operations are not uncommon. This illustrates GPFS' inherent trade‑off: token‑based strong consistency and concurrency control come at the cost of increased system complexity.&lt;/p&gt;

&lt;p&gt;Early InfiniBand networks or high‑quality fiber networks could reliably support this mechanism due to their low latency and high reliability. However, in some newer deployments—particularly those using RoCE—suboptimal network conditions, hardware quality, or operational expertise can amplify stability challenges in token coordination, especially when hundreds of clients coexist in a single cluster.  &lt;/p&gt;

&lt;p&gt;Overall, this mechanism enables high‑performance parallel access but demands high network, disk stability, and cluster operational capabilities. Client anomalies also require handling token recovery and cluster state restoration. Therefore, it’s essential to assess whether your team has the necessary deployment, monitoring, and troubleshooting skills.&lt;/p&gt;

&lt;p&gt;Based on practical experience, GPFS deployments also require attention to several engineering considerations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;mmap scenarios: GPFS employs special mechanisms for mmap, such as pagepool, to reduce memory copying and improve performance. However, this also creates complex interactions with OS memory management. Large‑scale reliance on mmap patterns is generally not recommended without thorough validation.
&lt;/li&gt;
&lt;li&gt;Hot files and large directories: Hot files, hot directories, massive small files, or many clients concurrently accessing the same directory can become performance bottlenecks due to metadata coordination and lock contention. Mitigation typically involves directory splitting, data sharding, and access pattern optimization.
&lt;/li&gt;
&lt;li&gt;Operational management and monitoring: GPFS' management interface is not particularly user‑friendly for newcomers, and some monitoring information is not intuitive. Many teams supplement with external monitoring systems like Grafana to better observe cluster state, performance metrics, and anomalies.
&lt;/li&gt;
&lt;li&gt;Cluster Export Services (CES), Active File Management (AFM), and other components: While these features enable more complex scenarios, they also introduce additional configuration, operational, and troubleshooting overhead. Teams without long‑term GPFS experience should carefully evaluate this complexity upfront.
&lt;/li&gt;
&lt;li&gt;Capacity planning: Many GPFS or CPFS deployments emphasize expansion over contraction. Capacity planning early on is crucial to avoid either excessive initial capacity (and long‑term cost) or insufficient capacity (hampering business growth).&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Performance comparison
&lt;/h3&gt;

&lt;p&gt;Intuitively, many people assume that JuiceFS—built on object storage and independent metadata services—cannot be directly compared to GPFS. However, in certain AI workloads, both systems do present comparable alternatives.  &lt;/p&gt;

&lt;p&gt;It’s important to note that GPFS performance depends heavily on the synergy of underlying disks, storage servers, networking, clients, and the parallel file system itself. JuiceFS, on the other hand, is influenced by object storage performance, metadata services, client caching, distributed cache groups, and mount modes. Therefore, comparisons should not be reduced to single performance numbers—they must account for I/O patterns, deployment architecture, data scale, and access paths.  &lt;/p&gt;

&lt;p&gt;The following tests were performed with &lt;a href="https://juicefs.com/docs/cloud/" rel="noopener noreferrer"&gt;JuiceFS Enterprise Edition&lt;/a&gt;. The Community Edition shares the same core architecture, so Community Edition users can also refer to these testing methodologies and results.&lt;/p&gt;

&lt;h4&gt;
  
  
  Sequential reads: GPFS better per node, JuiceFS scales through cache groups
&lt;/h4&gt;

&lt;p&gt;In single‑node scenarios, the performance models differ markedly. In our tests, with two 400 Gbps NICs, a single GPFS node achieved about 100 GB/s for sequential reads.  &lt;/p&gt;

&lt;p&gt;Under TCP mode (200 Gbps NIC), a single JuiceFS node peaked at about 20 GB/s. With RDMA (two 400 Gbps NICs), it reached about 55 GB/s. For higher aggregate throughput, JuiceFS can scale horizontally by adding cache nodes. For example, in one autonomous driving customer deployment with about 150 cache nodes (each with 160 Gbps NICs), the aggregated application throughput reached about 2.3 TB/s.&lt;/p&gt;

&lt;h4&gt;
  
  
  Sequential writes: GPFS excels at synchronous writes; JuiceFS uses writeback for scale
&lt;/h4&gt;

&lt;p&gt;For sequential writes, GPFS holds an advantage under synchronous write semantics. Its write capability stems from the underlying parallel storage system—data written is immediately accessible by other clients under strong consistency semantics. This is well suited for scenarios requiring write reliability, real‑time visibility, and consistency.&lt;/p&gt;

&lt;p&gt;JuiceFS sequential write performance depends on whether writeback is enabled. In synchronous mode, writes go to the backend object store, and performance is constrained by backend storage, protocol overhead, and network latency. With writeback enabled, data is first written to client‑local cache and later uploaded asynchronously—aggregate throughput improves, but real‑time visibility and consistency semantics change.  &lt;/p&gt;

&lt;p&gt;Thus, GPFS is better for scenarios requiring synchronous persistence and real‑time visibility. JuiceFS with writeback is suitable for workloads that can tolerate asynchronous upload semantics.&lt;/p&gt;

&lt;h4&gt;
  
  
  Random reads: GPFS better at high concurrency, JuiceFS competitive in certain scenarios
&lt;/h4&gt;

&lt;p&gt;In 4K single‑process random read tests, we compared JuiceFS, GPFS, and local disk /tmp (EXT4).  &lt;/p&gt;

&lt;p&gt;The figure below shows the test results:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;At iodepth 1, 2, and 4, JuiceFS outperformed GPFS.
&lt;/li&gt;
&lt;li&gt;At iodepth 8, GPFS surpassed JuiceFS and stabilized around 80K IOPS.
&lt;/li&gt;
&lt;li&gt;JuiceFS peaked around 68K IOPS at iodepth 4 and 8 and then gradually declined at deeper I/O depths.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffszi7wf3ropmljxeau4z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffszi7wf3ropmljxeau4z.png" alt=" " width="800" height="573"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Multi‑process random reads present more complex behaviors. When multiple processes concurrently read the same file, &lt;strong&gt;GPFS' consistency and lock mechanisms can introduce extra overhead&lt;/strong&gt;. To validate this, we tested both reading the same file and different files.  &lt;/p&gt;

&lt;p&gt;The following figure shows the test results:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GPFS reads of different files scaled rapidly with numjobs, reaching ~433K IOPS at numjobs=12.
&lt;/li&gt;
&lt;li&gt;GPFS reads of the same file dropped after numjobs=2, suggesting consistency coordination overhead.
&lt;/li&gt;
&lt;li&gt;JuiceFS reached ~258K IOPS at numjobs=12 and stabilized around 250K thereafter.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F41fbj31ou60ttmjorxag.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F41fbj31ou60ttmjorxag.png" alt=" " width="800" height="577"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;All tests were run with local caching disabled, and data was served from distributed cache. The results show that GPFS has higher random read performance under high concurrency, while JuiceFS remains highly capable—sufficient for most AI training requirements.&lt;/p&gt;

&lt;h4&gt;
  
  
  Random writes: GPFS superior at high concurrency
&lt;/h4&gt;

&lt;p&gt;Random writes better expose architectural differences between the two systems. JuiceFS' random write performance depends on whether writeback is enabled. Without it, performance is mainly constrained by the object storage backend. With writeback enabled, performance reflects the client's local cache path capabilities.  &lt;/p&gt;

&lt;p&gt;Under writeback enabled and &lt;code&gt;cache‑dir&lt;/code&gt; set to local NVMe storage, 4K multi‑process random write results showed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;From numjobs 1 to 3, JuiceFS increased from ~28K to ~56K IOPS, significantly higher than GPFS.
&lt;/li&gt;
&lt;li&gt;At numjobs=12, both were close (~51K vs. ~50K).
&lt;/li&gt;
&lt;li&gt;Beyond numjobs=16, GPFS continued to rise to ~66K (at 16) and ~81K (at 20), while JuiceFS was in the 50–60K range.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7k09e28ad5rm7v5vevs6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7k09e28ad5rm7v5vevs6.png" alt=" " width="800" height="590"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;These results indicate GPFS holds an advantage in random write performance, while JuiceFS closes the gap significantly when writeback is enabled. However, random writes are not common in AI workloads, so this should not be a primary evaluation criterion.&lt;/p&gt;

&lt;h4&gt;
  
  
  Selection summary
&lt;/h4&gt;

&lt;p&gt;The main differences between GPFS and JuiceFS are summarized in the table below: &lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;GPFS (symmetric, decentralized)&lt;/th&gt;
&lt;th&gt;JuiceFS (separated metadata &amp;amp; data)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Metadata architecture&lt;/td&gt;
&lt;td&gt;Distributed, embedded in local memory&lt;/td&gt;
&lt;td&gt;Independent, high-performance external database (cloud-native, lightweight)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data storage layer&lt;/td&gt;
&lt;td&gt;Expensive, tightly coupled shared SAN / parallel disk&lt;/td&gt;
&lt;td&gt;Cost-effective, highly reliable, highly elastic object storage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Locking &amp;amp; concurrency&lt;/td&gt;
&lt;td&gt;Distributed token-based locks (strong consistency)&lt;/td&gt;
&lt;td&gt;Optimistic concurrency (Community Edition) / Single-threaded core (Enterprise Edition)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical use cases&lt;/td&gt;
&lt;td&gt;Select HPC and scientific computing scenarios&lt;/td&gt;
&lt;td&gt;AI research, training, inference acceleration, large-scale data management&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;GPFS is better suited for environments with ample budgets and stringent requirements for low latency, strong consistency, high‑concurrency access, and random write capability—typical of traditional HPC, scientific computing, and some quantitative finance applications. Its performance also depends on stable networking, storage hardware, and professional operational expertise.  &lt;/p&gt;

&lt;p&gt;For AI research, training, inference acceleration, and large‑scale data management challenges involving scalability, performance, cost, and multi‑cloud management, JuiceFS is the more suitable solution—and has been validated in production across all the aforementioned domains.&lt;/p&gt;

&lt;h2&gt;
  
  
  Alluxio vs. JuiceFS
&lt;/h2&gt;

&lt;p&gt;Alluxio is another solution that enterprises often compare with JuiceFS when evaluating AI storage architectures.&lt;/p&gt;

&lt;p&gt;Both solutions can provide filesystem access and cache acceleration on top of object storage. However, they differ significantly in product positioning, data organization models, and metadata architectures.&lt;/p&gt;

&lt;p&gt;From the perspective of product evolution, the two projects have taken different approaches:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Alluxio’s recent AI-focused development has mainly concentrated on its Enterprise AI product line. The latest release of the Alluxio open-source repository is v2.9.4, released in June 2024.
&lt;/li&gt;
&lt;li&gt;JuiceFS continues to evolve both its open-source and enterprise editions. Many new capabilities are first introduced, validated, and refined in the open-source edition before gradually becoming available in the enterprise edition to serve more production users.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Key architectural differences
&lt;/h3&gt;

&lt;h4&gt;
  
  
  Data organization and consistency boundaries
&lt;/h4&gt;

&lt;p&gt;Alluxio uses a &lt;strong&gt;1:1 transparent caching model&lt;/strong&gt;, which preserves the original organization of files in the underlying storage system. Existing data does not need to be imported or reorganized in advance, and the caching layer can be added or removed as needed.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx62xy6nfhnwqd1cbxgdl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx62xy6nfhnwqd1cbxgdl.png" alt=" " width="800" height="581"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;JuiceFS, on the other hand, splits files into data chunks and stores them in object storage. It uses a metadata service to maintain file system semantics and data chunk mappings. From the application perspective, users see a complete file system, while the objects stored in the backend object storage are JuiceFS-managed data chunks rather than the original files.&lt;/p&gt;

&lt;p&gt;This difference also affects the &lt;strong&gt;source of truth&lt;/strong&gt; and consistency boundaries:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;For Alluxio, the actual data typically remains in object storage or other UFS systems. The cache layer mainly accelerates access. If applications bypass Alluxio and directly modify data in the underlying storage, additional cache invalidation, refresh, or synchronization mechanisms may be required.
&lt;/li&gt;
&lt;li&gt;JuiceFS combines metadata services and object storage into a complete file system. File states and data mappings are managed centrally by JuiceFS, and application reads and writes go through JuiceFS. Therefore, its consistency boundary exists within the file system itself and does not rely on additional synchronization between the cache layer and backend storage.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Cache and namespace organization
&lt;/h3&gt;

&lt;p&gt;Alluxio emphasizes a unified namespace and shared cache pool. It can integrate multiple underlying storage systems, including OSS, S3, HDFS, Ceph, MinIO, and NAS, into a single namespace and accelerate access through a unified distributed cache.&lt;/p&gt;

&lt;p&gt;For organizations that already have multiple storage systems and do not want to migrate or reorganize existing data, this approach provides greater flexibility. However, when multiple workloads share the same cache pool, resource management and isolation typically require policies such as directory-based controls, priorities, and time to live (TTL) settings.&lt;/p&gt;

&lt;p&gt;JuiceFS typically uses an independent filesystem as the management unit. Different object storage backends, such as OSS, COS, and TOS, usually correspond to different file systems with their own cache configurations.&lt;/p&gt;

&lt;p&gt;Compared with Alluxio’s emphasis on a unified view across multiple data sources and shared caching, &lt;strong&gt;JuiceFS focuses more on clear boundaries between file system management, permissions, and data governance&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;In addition, JuiceFS Enterprise Edition can connect multiple buckets into a single file system and use the same cache resources for acceleration. This enables multi-source data access scenarios while maintaining clear file system-level management boundaries.&lt;/p&gt;

&lt;h3&gt;
  
  
  Metadata architecture
&lt;/h3&gt;

&lt;p&gt;Alluxio Enterprise distributes caching and part of the state management responsibilities across workers. Workers not only provide data caching but also participate in managing cache status and data location information. This brings data access closer to compute nodes or GPU nodes. Coordination and management components still exist, but cache-related states are not entirely centralized in an independent metadata service. As a result, when workers frequently restart or scale in and out, additional attention is required for recovering cache states and coordinating data locations.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv1jr42v4oxnxby1t1fnh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv1jr42v4oxnxby1t1fnh.png" alt=" " width="800" height="495"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;JuiceFS manages core file system metadata through an independent metadata service.&lt;/p&gt;

&lt;p&gt;The open-source edition supports metadata engines such as Redis and TiKV, while the enterprise edition provides a highly available metadata service to support consistency and transactional operations. Distributed cache nodes are responsible only for caching data blocks. Therefore, cache node failures mainly affect cache hit rates and access performance, but they do not change file system metadata states.&lt;/p&gt;

&lt;h3&gt;
  
  
  How architectural differences affect real-world usage
&lt;/h3&gt;

&lt;h4&gt;
  
  
  POSIX compatibility
&lt;/h4&gt;

&lt;p&gt;Alluxio exposes data from object storage or other UFS systems through file system interfaces, but its file system semantics are not always complete.  &lt;/p&gt;

&lt;p&gt;Capabilities such as timestamps, file locking, hard links, symbolic links, extended attributes, and ACLs may require additional configuration or have usage limitations. For read acceleration and lightweight access scenarios, these limitations may not be significant.&lt;/p&gt;

&lt;p&gt;However, if an application intends to use Alluxio as a full-featured file system, these factors should be carefully evaluated.&lt;/p&gt;

&lt;p&gt;JuiceFS aims to provide complete file system capabilities on top of object storage. Therefore, JuiceFS continuously improves POSIX compatibility, covering common file system semantics such as directories, permissions, timestamps, and file locks, as well as more advanced system calls such as ioctl-based immutable and read-only settings. In production environments, JuiceFS POSIX compatibility already covers the vast majority of workloads and approaches full POSIX compatibility.&lt;/p&gt;

&lt;h4&gt;
  
  
  Write performance and write amplification
&lt;/h4&gt;

&lt;p&gt;Alluxio preserves the original file format, which provides transparent access and reduces the cost of onboarding existing datasets. However, random writes, overwrites, and partial updates require careful consideration of write amplification. The reason is that object storage generally does not support in-place modification of object contents. If the backend maintains a 1:1 file layout, a small modification may trigger rewriting a larger amount of data, potentially requiring the entire object or file to be rewritten.&lt;/p&gt;

&lt;p&gt;JuiceFS uses a chunk-based data layout. For random writes or append operations, JuiceFS only needs to process affected data chunks and update corresponding metadata, without being constrained by the original object storage file layout. Therefore, JuiceFS can better control write amplification for partial updates and random write workloads. However, if users need to restore JuiceFS data into the original file format in object storage, additional export or migration operations are required.&lt;/p&gt;

&lt;h4&gt;
  
  
  Deployment and engineering capabilities
&lt;/h4&gt;

&lt;p&gt;Alluxio Enterprise supports full-component Kubernetes deployment. It also supports temporary write caching without persisting data to object storage, such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;temporary decompression workloads
&lt;/li&gt;
&lt;li&gt;intermediate computation results
&lt;/li&gt;
&lt;li&gt;short-lived data that can be discarded after use&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These capabilities are closer to cache-layer acceleration and compute-side optimization scenarios.  &lt;/p&gt;

&lt;p&gt;JuiceFS Enterprise Edition currently typically deploys metadata services on virtual machines or physical servers, mainly due to requirements for metadata stability and system reliability.  &lt;/p&gt;

&lt;p&gt;JuiceFS focuses more on long-term filesystem capabilities, including:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;large-scale metadata management
&lt;/li&gt;
&lt;li&gt;Trash and data recovery
&lt;/li&gt;
&lt;li&gt;transactional operations
&lt;/li&gt;
&lt;li&gt;cache governance
&lt;/li&gt;
&lt;li&gt;seamless upgrades
&lt;/li&gt;
&lt;li&gt;open-source ecosystem integration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In production, JuiceFS has already supported up to 500 billion files—an important demonstration of its capability as a complete file system in large‑scale data management scenarios.&lt;/p&gt;

&lt;h3&gt;
  
  
  Selection recommendations
&lt;/h3&gt;

&lt;p&gt;The difference between Alluxio and JuiceFS is not simply about the number of features. They are designed to solve different problems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If data already exists in object storage, HDFS, Ceph, MinIO, or NAS, and organizations do not want to migrate or reorganize existing data but only need a pluggable cache layer to accelerate reads, Alluxio can be a suitable choice.
&lt;/li&gt;
&lt;li&gt;JuiceFS is likely a better fit, if organizations want to build a complete file system on top of object storage while supporting:

&lt;ul&gt;
&lt;li&gt;mixed read/write workloads
&lt;/li&gt;
&lt;li&gt;full POSIX semantics
&lt;/li&gt;
&lt;li&gt;strong consistency
&lt;/li&gt;
&lt;li&gt;elastic metadata scaling
&lt;/li&gt;
&lt;li&gt;tens of thousands of concurrent clients
&lt;/li&gt;
&lt;li&gt;multi-cloud data management&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;GPFS, Alluxio, and JuiceFS address different core storage challenges:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GPFS&lt;/strong&gt; focuses on low latency and high-concurrency read/write performance for high-performance computing workloads.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Alluxio&lt;/strong&gt; primarily provides cache acceleration for existing datasets.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;JuiceFS&lt;/strong&gt; provides a complete, scalable file system built on top of object storage.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Therefore, storage selection should not be based only on individual performance metrics. Organizations should first determine whether they need high-performance shared storage, a data acceleration layer, or a complete file system designed for multi-cloud and large-scale data management.&lt;/p&gt;

&lt;p&gt;The final decision should consider consistency requirements, data scale, cost, and operational complexity.&lt;/p&gt;

&lt;p&gt;If you have any questions for this article, feel free to join &lt;a href="https://github.com/juicedata/juicefs/discussions/" rel="noopener noreferrer"&gt;JuiceFS discussions on GitHub&lt;/a&gt; and &lt;a href="http://go.juicefs.com/discord" rel="noopener noreferrer"&gt;the community on Discord&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>5.5 Faster Small-File Writes: Tuhu Built a Unified AI Storage Platform with JuiceFS + Ceph RADOS</title>
      <dc:creator>DASWU</dc:creator>
      <pubDate>Tue, 04 Aug 2026 02:46:16 +0000</pubDate>
      <link>https://dev.to/daswu/55x-faster-small-file-writes-tuhu-built-a-unified-ai-storage-platform-with-juicefs-ceph-rados-5e1l</link>
      <guid>https://dev.to/daswu/55x-faster-small-file-writes-tuhu-built-a-unified-ai-storage-platform-with-juicefs-ceph-rados-5e1l</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;em&gt;TL;DR:&lt;/em&gt;&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;em&gt;Tuhu leverages JuiceFS, TiKV, and Ceph RADOS to unify storage for AI training, AI inference, and analytics, supporting over 100 million files while achieving up to **5.5× faster small-file writes&lt;/em&gt;* and &lt;strong&gt;3.2× faster small-file reads&lt;/strong&gt;.*&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.tuhu.cn/us/home/tuhuaboutus?click=firstMenu" rel="noopener noreferrer"&gt;Tuhu&lt;/a&gt; (9690.HK) is an integrated online and offline automotive service platform. By the end of 2025, we had grown to 162.3 million registered users and operated 8,008 service centers across China. As our business continued to expand, our infrastructure had to support increasingly diverse data management and compute workloads. Over the years, multiple storage systems—including NFS, Alluxio, MinIO, and SeaweedFS—had been deployed independently to meet different application requirements. While each solution addressed a specific use case, the overall storage architecture became fragmented, making data movement expensive and operations increasingly difficult.&lt;/p&gt;

&lt;p&gt;When evaluating a unified storage architecture, we built a unified &lt;a href="https://www.ibm.com/think/topics/ai-storage" rel="noopener noreferrer"&gt;AI storage&lt;/a&gt; platform based on &lt;strong&gt;JuiceFS + TiKV + Ceph object storage cluster (RADOS)&lt;/strong&gt;, considering our existing private cloud infrastructure and Ceph deployment.&lt;/p&gt;

&lt;p&gt;Today, this platform serves &lt;a href="https://www.akamai.com/glossary/what-is-ai-training" rel="noopener noreferrer"&gt;AI training&lt;/a&gt;, &lt;a href="https://www.ibm.com/think/topics/ai-inference" rel="noopener noreferrer"&gt;AI inference&lt;/a&gt;, and big data processing workloads through a single storage infrastructure. It currently manages over 100 million files. In low-concurrency benchmark tests, &lt;strong&gt;it achieved up to a 5.5x improvement in small-file sequential write performance and up to a 3.2x improvement in small-file read performance&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;In this article, we’ll explain how we designed and deployed our unified AI storage platform on top of its existing Ceph infrastructure. We’ll also share several production optimization practices, including Ceph RADOS data paths, erasure-coded pools, write amplification reduction for small files, and stability improvements for containerized deployments.&lt;/p&gt;

&lt;h2&gt;
  
  
  From multiple storage systems to a unified storage foundation
&lt;/h2&gt;

&lt;p&gt;In the early stages of our company's growth, storage infrastructure was built independently for different applications to maximize development speed and satisfy immediate application requirements. As the platform evolved, this approach gradually resulted in multiple storage systems coexisting in production.&lt;/p&gt;

&lt;p&gt;The rapid adoption of cloud-native computing, large-scale AI training, and AI inference exposed the limitations of this architecture. Problems related to architectural consistency, data mobility, operational complexity, and storage performance became increasingly apparent.&lt;/p&gt;

&lt;p&gt;We identified four major challenges.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;High operational complexity:&lt;/strong&gt; Production environments relied on numerous storage solutions, including NFS, Alluxio, MinIO, Ceph, SeaweedFS, and various cloud block storage services. Each system had its own deployment model, access protocol, operational workflow, and troubleshooting methodology. As storage clusters continued to grow, day-to-day maintenance, capacity planning, and version upgrades became increasingly difficult.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://www.ibm.com/think/topics/data-silos" rel="noopener noreferrer"&gt;&lt;strong&gt;Data silos&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;:&lt;/strong&gt; Different storage systems exposed different interfaces and lacked a unified data access layer. AI training pipelines typically require POSIX-compliant file semantics, while surrounding tools for data processing, model management, backup, and archiving often depend on object storage APIs. Because these storage systems could not naturally share the same data, datasets frequently had to be copied or synchronized between platforms. This increased operational overhead, storage consumption, and workflow complexity.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Rising storage costs:&lt;/strong&gt; As business continued to grow, both the number of files and total storage capacity increased rapidly. Meanwhile, infrastructure costs—including hardware procurement, datacenter resources, and storage operations—continued to rise. Improving storage utilization while maintaining performance and reliability became a key objective for the infrastructure team.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Higher performance requirements from AI workloads:&lt;/strong&gt; AI training, AI inference, disaggregated storage and compute, and big data analytics placed higher demands on the underlying storage system's concurrency, throughput, access latency, and stability. In particular, large-scale small-file access, multi-node concurrent reads, model file loading, and online inference services required the storage system not only to provide stable capacity support but also to deliver predictable performance under high-concurrency conditions.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As a result, the existing architecture based on multiple independent storage systems could no longer support long-term application growth. We decided to consolidate fragmented storage resources into a unified storage foundation that would simplify operations while enabling seamless data sharing across different compute environments.&lt;/p&gt;

&lt;p&gt;The new storage platform was designed around two major capability areas:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Block storage for infrastructure and virtualization workloads&lt;/strong&gt;: This capability provided reliable, high-performance storage for virtual machine images, cloud disks, and traditional infrastructure services.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Unified data access for &lt;a href="https://cloud.google.com/learn/what-is-cloud-native" rel="noopener noreferrer"&gt;cloud-native&lt;/a&gt; workloads&lt;/strong&gt;, including Kubernetes applications, traditional microservices, big data analytics, AI training, AI inference, and emerging middleware such as vector databases. Among these, AI training and inference posed more concentrated challenges in terms of file semantics, object interfaces, data throughput, small-file performance, and multi-node concurrent access.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Therefore, building a unified cloud storage foundation involved both block storage capability for virtualized environments and data access capability for AI and cloud-native workloads. The following section will focus on the latter—how to build a unified data access foundation based on JuiceFS to support AI training, AI inference, and big data processing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building a unified AI storage platform with JuiceFS
&lt;/h2&gt;

&lt;p&gt;When designing a unified storage platform, we evaluated several architectural approaches based on our existing private cloud infrastructure, storage investments, and workload characteristics.&lt;/p&gt;

&lt;p&gt;Our primary goals were to support large-scale data access, enable disaggregated storage and compute, provide multiple access protocols, and deliver high reliability without introducing unnecessary operational complexity.&lt;/p&gt;

&lt;p&gt;Rather than building a new distributed storage system from scratch, we chose a modular architecture based on &lt;strong&gt;JuiceFS, TiKV, and Ceph RADOS&lt;/strong&gt;. This approach allowed us to leverage our existing Ceph infrastructure while introducing a unified file system capable of serving AI training, inference, and big data processing workloads through a single storage platform.&lt;/p&gt;

&lt;p&gt;A key design principle of this architecture is the separation of metadata and user data.&lt;/p&gt;

&lt;p&gt;JuiceFS provides the unified file system layer, while the metadata engine and Ceph RADOS independently handle metadata management and data storage. At the storage layer, Ceph distributes data and manages failure domains using the CRUSH algorithm, providing a decentralized and highly reliable storage backend.&lt;/p&gt;

&lt;p&gt;The architecture also adopts a &lt;strong&gt;lightweight client&lt;/strong&gt; model.&lt;/p&gt;

&lt;p&gt;JuiceFS clients are responsible for file system semantics, metadata operations, cache management, and request scheduling. Once data is written to Ceph RADOS, data durability—whether implemented through replication or erasure coding—is handled entirely by the storage cluster.&lt;/p&gt;

&lt;p&gt;This division of responsibilities keeps compute nodes lightweight, minimizes interference with AI training workloads, and fully utilizes the reliability and scalability already provided by the existing Ceph deployment.&lt;/p&gt;

&lt;p&gt;For our private cloud environment, this modular approach significantly reduced engineering and operational overhead compared with developing and maintaining an integrated distributed storage system. At the same time, it provides unified access to AI training, AI inference, and big data workloads through a common storage platform.&lt;/p&gt;

&lt;h3&gt;
  
  
  Unified AI storage architecture
&lt;/h3&gt;

&lt;p&gt;Our unified AI storage platform consists of four logical layers.&lt;/p&gt;

&lt;h4&gt;
  
  
  Workload layer
&lt;/h4&gt;

&lt;p&gt;At the top are the application workloads, including virtual machines, Kubernetes-based containerized applications, traditional microservices, AI training, AI inference, and Apache Spark jobs.&lt;/p&gt;

&lt;h4&gt;
  
  
  Cache layer
&lt;/h4&gt;

&lt;p&gt;The cache layer consists of two components:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;DataCache Pool:&lt;/strong&gt; It accelerates repeated access to training datasets, model files, and other hot data, reducing load on the storage backend.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;KVCache Pool&lt;/strong&gt;: It’s designed for &lt;a href="https://en.wikipedia.org/wiki/Large_language_model" rel="noopener noreferrer"&gt;large language model&lt;/a&gt; (LLM) inference. It caches context data across multiple storage tiers, reducing latency and improving the responsiveness of online inference services.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9p3w1iptjw0jwtsnrgui.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9p3w1iptjw0jwtsnrgui.png" alt=" " width="800" height="368"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h4&gt;
  
  
  Storage protocol layer
&lt;/h4&gt;

&lt;p&gt;JuiceFS provides a unified data access entry point here. AI training jobs access data via POSIX FUSE or CSI Driver; big data Spark jobs access data through the JuiceFS Hadoop Java SDK; and object storage toolchains access the same data via &lt;a href="https://juicefs.com/docs/community/guide/gateway" rel="noopener noreferrer"&gt;JuiceFS S3 Gateway&lt;/a&gt;. This approach unifies file storage, object storage, and big data access interfaces around a single underlying data set, reducing data movement and redundant storage across systems.&lt;/p&gt;

&lt;p&gt;It should be noted that block storage scenarios, such as VM cloud disks, are primarily handled by Ceph RBD, which is part of the overall cloud storage system's block storage capability. The focus of this article is the file, object, and big data access foundation built on JuiceFS. While both share the underlying Ceph storage resources, they differ in access semantics and service targets.&lt;/p&gt;

&lt;h4&gt;
  
  
  Storage engine layer
&lt;/h4&gt;

&lt;p&gt;JuiceFS uses a &lt;a href="https://juicefs.com/docs/community/architecture/" rel="noopener noreferrer"&gt;metadata-data decoupled architecture&lt;/a&gt;. For file systems with hundreds of millions of files, metadata capabilities directly impact path resolution, directory traversal, file attribute queries, small-file access, and concurrent job startup performance. TiKV provides distributed transactions, strong consistency, high availability, and horizontal scalability, offering stable metadata support for large-scale file systems. We built a five-node TiKV cluster for the metadata layer.&lt;/p&gt;

&lt;p&gt;For the data layer, we use Ceph RADOS as the underlying data storage engine for JuiceFS. Ceph RADOS reuses existing private cloud storage resources, reducing duplicate infrastructure costs. In addition, by creating an erasure-coded pool for physical data storage, we improve physical disk capacity utilization while maintaining reliability, mitigating cost pressures from data growth.&lt;/p&gt;

&lt;p&gt;In this architecture, data communication follows two paths:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;The metadata path between clients and TiKV, handling path lookups, directory structures, file attributes, and state information&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The data read/write path between clients and Ceph RADOS, handling actual data block access&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Through librados, JuiceFS interacts directly with the underlying Ceph RADOS. Compared to traditional gateway-based storage architectures, this reduces multi-layer forwarding overhead, allowing clients to access the storage cluster more directly after metadata interactions and shortening the data access path.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi7vpe7b082c9sd0kka2v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi7vpe7b082c9sd0kka2v.png" alt=" " width="799" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Production deployment and optimization practices
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Critical scenario deployment and storage consolidation
&lt;/h3&gt;

&lt;p&gt;In LLM application scenarios, our company has dedicated self-built compute clusters for model inference services such as MiniMax. Meanwhile, custom models developed by internal application lines have gradually been migrated to JuiceFS after passing production validation. As a result, core LLM inference and training applications have begun to unify onto the new file storage foundation, with the original NFS clusters and Alluxio storage clusters entering phased decommissioning.&lt;/p&gt;

&lt;p&gt;In core middleware and backup scenarios, including large-scale ClickHouse deployments, multiple core middleware systems and full production backup tasks have also been migrated to JuiceFS. As the unified file foundation has proven its capacity, stability, and access efficiency, MinIO and SeaweedFS have also entered decommissioning, reducing the operational complexity of maintaining multiple parallel storage systems.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cache acceleration for AI training and inference
&lt;/h3&gt;

&lt;p&gt;For AI training scenarios, training data typically has a "write once, read many" pattern. We use JuiceFS Community Edition's local data caching capability to cache hot training data on compute nodes. This reduces backend storage pressure and improves multi-node concurrent read efficiency.&lt;/p&gt;

&lt;p&gt;For LLM inference scenarios, models serving online requests or performing iterative token generation are sensitive to context data loading latency and throughput. With limited single-node GPU memory and memory capacity, we use a multi-level caching mechanism to cache context data across multiple levels, improving online inference response efficiency.&lt;/p&gt;

&lt;h3&gt;
  
  
  Benchmarking and small I/O optimization
&lt;/h3&gt;

&lt;p&gt;After deploying the unified storage platform, we focused on evaluating its performance under small-file workloads.&lt;/p&gt;

&lt;p&gt;Using &lt;code&gt;juicefs bench&lt;/code&gt; and fio, we benchmarked small-file operations and random I/O to compare the new JuiceFS + Ceph RADOS architecture with the previous JuiceFS + MinIO deployment.&lt;/p&gt;

&lt;p&gt;To minimize the impact of caching, all local caches were disabled during testing, allowing I/O requests to be served directly by the storage backend. It's worth noting that these benchmarks were intended to validate the architecture rather than measure peak cluster performance. The tests were conducted with only four client threads under low concurrency.&lt;/p&gt;

&lt;p&gt;Even under low-concurrency test conditions, the new architecture demonstrated substantial performance improvements.&lt;/p&gt;

&lt;p&gt;The two tables below show test results for JuiceFS + MinIO:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Throughput&lt;/th&gt;
&lt;th&gt;Latency&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Large file write&lt;/td&gt;
&lt;td&gt;3,377.65 MiB/s&lt;/td&gt;
&lt;td&gt;1.21 s/file&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Large file read&lt;/td&gt;
&lt;td&gt;4,248.77 MiB/s&lt;/td&gt;
&lt;td&gt;0.96 s/file&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Small file write&lt;/td&gt;
&lt;td&gt;147.2 files/s&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;27.17&lt;/strong&gt; ms/file&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Small file read&lt;/td&gt;
&lt;td&gt;1,176.8 files/s&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;3.40&lt;/strong&gt; ms/file&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;IOPS&lt;/th&gt;
&lt;th&gt;Throughput (MiB/s)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Random read&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;57,504&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1,483&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Random write&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1,194&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1,521&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two tables below show test results for JuiceFS + Ceph RADOS:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Throughput&lt;/th&gt;
&lt;th&gt;Latency&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Large file write&lt;/td&gt;
&lt;td&gt;4,984.12 MiB/s&lt;/td&gt;
&lt;td&gt;0.82 s/file&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Large file read&lt;/td&gt;
&lt;td&gt;4,817.87 MiB/s&lt;/td&gt;
&lt;td&gt;0.85 s/file&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Small file write&lt;/td&gt;
&lt;td&gt;817.0 files/s&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;4.90&lt;/strong&gt; ms/file&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Small file read&lt;/td&gt;
&lt;td&gt;3,776.4 files/s&lt;/td&gt;
&lt;td&gt;
&lt;strong&gt;1.06&lt;/strong&gt; ms/file&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;IOPS&lt;/th&gt;
&lt;th&gt;Throughput (MiB/s)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Random read&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;77,984&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;3,394&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Random write&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3,836&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;4,120&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Test result highlights:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Small-file write throughput improved by up to &lt;strong&gt;5.5x&lt;/strong&gt;, with significantly lower latency for individual file writes.&lt;/li&gt;
&lt;li&gt;Small-file read throughput improved by up to &lt;strong&gt;3.2x&lt;/strong&gt;, delivering higher read throughput and lower per-file read latency.&lt;/li&gt;
&lt;li&gt;Random I/O performance also improved. Both throughput and IOPS increased, with random-write IOPS improving by more than &lt;strong&gt;3x&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;From an architectural perspective, the JuiceFS + MinIO solution requires S3 interface and object storage gateway overhead, while the JuiceFS + Ceph RADOS solution uses librados for more direct interaction with the underlying storage cluster. This reduces gateway forwarding overhead. As a result, &lt;strong&gt;Ceph RADOS shows more significant performance advantages in scenarios sensitive to access paths, such as small-file writes and random writes&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;However, in random small-file read tests, since local caching was disabled, data still needed to be retrieved from physical disks in real time. Therefore, read performance did not show the same order-of-magnitude improvement as writes, but still achieved a stable ~1.4x improvement overall.&lt;/p&gt;

&lt;h3&gt;
  
  
  Small-file space amplification and write amplification optimization
&lt;/h3&gt;

&lt;p&gt;In erasure coding (EC) mode, traditional EC algorithms require alignment with stripe units. In a 6+3 configuration, even writing a 10-byte tiny file requires filling 6 data chunks to 4 KB each plus 3 coding chunks, resulting in 36 KB of physical storage—causing severe write amplification and thousands of times space waste.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flo4dznq3477d1373cvfb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flo4dznq3477d1373cvfb.png" alt=" " width="799" height="297"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;To address this, we introduced &lt;a href="https://ceph.io/en/news/blog/2025/tentacle-fastec-performance-updates/" rel="noopener noreferrer"&gt;Ceph's Fast EC feature&lt;/a&gt; to optimize small-object writes.&lt;/p&gt;

&lt;p&gt;On one hand, Fast EC alleviates small-file space amplification. When enabled, small objects are no longer forced to fill entire stripes. For a 10-byte file, the logical size remains 10 bytes, while physical space occupies the first stripe; with a 4 KB minimum allocation unit and parity data, physical footprint is reduced to 16 KB.&lt;/p&gt;

&lt;p&gt;On the other hand, Fast EC also improves read/write efficiency in small I/O scenarios. Reads no longer require fetching full stripes; writes can use &lt;a href="https://ceph.io/en/news/blog/2025/tentacle-fastec-performance-updates/#parity-delta-writes" rel="noopener noreferrer"&gt;parity delta writes&lt;/a&gt; (PDW) to reduce unnecessary network interactions and I/O overhead, improving overall small-file read/write performance.&lt;/p&gt;

&lt;h3&gt;
  
  
  JuiceFS Mount Pod stability optimization
&lt;/h3&gt;

&lt;p&gt;In containerized environments, AI training applications occasionally encountered bad file descriptor errors, which severely impacted training efficiency. Investigation revealed that these errors were related to Mount Pod out-of-memory (OOM) conditions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Periodic OOM due to metadata backups:&lt;/strong&gt; When the cluster reached tens of millions or hundreds of millions of files, JuiceFS' default hourly metadata backups generated large backup files. Since these backup tasks were randomly assigned to clients in the cluster, and the underlying RADOS architecture imposes size limits on individual object writes (for example, 128 MB), backup tasks could fail repeatedly in large-scale file system scenarios, causing Mount Pod memory usage to grow continuously and eventually trigger OOM.&lt;/p&gt;

&lt;p&gt;To address this, we disabled automatic backup on Mount Pods, preventing online clients from performing automatic backups. Instead, dedicated scheduled tasks handle metadata backups centrally, avoiding additional backup pressure on application clients.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhoeauepw9n1ol48o0deg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhoeauepw9n1ol48o0deg.png" alt=" " width="566" height="209"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Occasional OOM due to page cache growth:&lt;/strong&gt; During frequent reads of large amounts of local cache, the Linux kernel page cache continued to grow. In such cases, even if the container's actual memory usage was only about 1.1GB, Kubernetes OOM decisions based on &lt;code&gt;working_set_bytes&lt;/code&gt;—which includes page cache—could trigger OOM kills when the metric rose to high levels (for example, 33.5 GB), causing occasional client disconnections.&lt;/p&gt;

&lt;p&gt;To address this, we increased &lt;code&gt;limits.memory&lt;/code&gt; based on historical monitoring data. Then, we physically isolated cache resources, ensuring that different PVCs did not share the same local cache directory. For nodes with limited memory (such as CPU nodes), we capped local cache capacity (data cache size) within a reasonable range (for example, 10 GB to 50 GB) and enabled cache reclamation once the limit was reached. This prevents uncontrolled growth of the kernel page cache inside the container.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjntxtub8qhvodxlk89z9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjntxtub8qhvodxlk89z9.png" alt=" " width="800" height="188"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Through these optimizations, Mount Pod stability improved significantly in large-scale file system and high-frequency cache access scenarios. This provides a solid foundation for onboarding more AI training and inference workloads to the unified file foundation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Future plans
&lt;/h2&gt;

&lt;p&gt;We plan to focus on three aspects for the continued evolution of the AI cloud storage infrastructure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Unified management from private cloud to &lt;a href="https://cloud.google.com/learn/what-is-hybrid-cloud" rel="noopener noreferrer"&gt;hybrid cloud&lt;/a&gt;:&lt;/strong&gt; As AI training jobs are scheduled across both private and public clouds, the storage foundation needs stronger cross-environment data management capabilities. Building on the existing private Ceph RADOS architecture, we’ll continue exploring JuiceFS' integration with public cloud object storage.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;By pushing cross-cloud data synchronization, data distribution, and access path management down to the storage layer as much as possible, we can reduce the application's awareness of differences between cloud environments, providing more consistent mount paths and access experiences. This will improve training job migration efficiency across different compute environments and reduce data management complexity in multi-cloud settings.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Promoting lakehouse integration to reduce cross-system data movement:&lt;/strong&gt; Previously, data movement between offline HDFS big data clusters and algorithm storage clusters relied on synchronization tools, increasing pipeline complexity, data redundancy, sync latency, and management costs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We plan to further adopt and refine the JuiceFS Hadoop Java SDK so that Spark jobs in the algorithm layer can directly read and write the underlying unified storage pool. This will allow big data computing, algorithm training, and data archiving to operate around a single storage foundation, reducing cross-system data replication and redundant storage, further breaking down data silos, and advancing storage-compute separation and lakehouse integration.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Building tiered storage to accelerate LLM inference&lt;/strong&gt;: As LLM inference services scale, a single storage medium can no longer simultaneously meet cost, capacity, and latency requirements. We plan to build capabilities in data tiering and KVCache tiering.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;For data tiering:&lt;/strong&gt; JuiceFS' current data tiering feature is primarily designed for public cloud object storage. For private Ceph RADOS environments, we plan to collaborate with the community to explore mounting multiple Ceph storage pools of different performance tiers (for example, NVMe Pool and HDD Pool) under the same JuiceFS file system, allowing data with different access frequencies to reside on appropriate storage tiers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For KVCache tiering:&lt;/strong&gt; As LLM inference evolves toward prefill/decode (PD) separation in distributed architectures, relying solely on single-node GPU memory to hold context data faces capacity and cost bottlenecks. Building on the unified storage foundation, we’ll gradually explore a four-tier KVCache architecture composed of HBM (GPU memory), DRAM, local SSD, and shared storage, providing more flexible context data management and access acceleration for large-scale AI inference clusters.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you have any questions for this article, feel free to join &lt;a href="https://github.com/juicedata/juicefs/discussions/" rel="noopener noreferrer"&gt;JuiceFS discussions on GitHub&lt;/a&gt; and &lt;a href="http://go.juicefs.com/discord" rel="noopener noreferrer"&gt;community on Discord&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Amazon S3 Files: How It Works, Performance Boundaries, and a Comparison with JuiceFS</title>
      <dc:creator>DASWU</dc:creator>
      <pubDate>Fri, 24 Jul 2026 02:45:46 +0000</pubDate>
      <link>https://dev.to/daswu/amazon-s3-files-how-it-works-performance-boundaries-and-a-comparison-with-juicefs-5fbi</link>
      <guid>https://dev.to/daswu/amazon-s3-files-how-it-works-performance-boundaries-and-a-comparison-with-juicefs-5fbi</guid>
      <description>&lt;p&gt;AWS announced &lt;a href="https://aws.amazon.com/s3/features/files/" rel="noopener noreferrer"&gt;Amazon S3 Files&lt;/a&gt; to let you mount an S3 bucket as a high‑performance shared file system on compute nodes without moving any data.  &lt;/p&gt;

&lt;p&gt;This isn't the industry's first attempt to make S3 accessible as a file system. From the early days of s3fs, to AWS' later release of &lt;a href="https://github.com/awslabs/mountpoint-s3" rel="noopener noreferrer"&gt;Mountpoint for Amazon S3&lt;/a&gt;, and now S3 Files, the path to "using S3 like a file system" has been many years in the making. The key difference is that earlier solutions mostly operated at the access layer, while S3 Files finally wraps shared access, full file system semantics, and a managed high‑performance layer into a native offering.  &lt;/p&gt;

&lt;p&gt;This makes S3 Files a new option worth analyzing on its own. For workloads that need file‑style access to existing S3 data, it provides a native, lightweight solution. But in more complex scenarios such as AI model training or big data analytics, its actual performance still needs to be evaluated based on its underlying implementation and runtime behavior.&lt;/p&gt;

&lt;p&gt;In this article, we’ll analyze S3 Files’ implementation, performance boundaries, and how it differs from JuiceFS.&lt;/p&gt;

&lt;h2&gt;
  
  
  S3 Files: an S3 native file system with EFS as the high‑performance layer
&lt;/h2&gt;

&lt;p&gt;S3 Files uses Amazon Elastic File System (EFS) as a managed high‑performance storage layer to handle low‑latency data and metadata. On top of that, it provides full file system semantics for S3, including consistency, file locking, and POSIX permissions.&lt;br&gt;&lt;br&gt;
Think of it this way: AWS adds an EFS-based file system access layer on top of object storage, so that data previously accessible only through object APIs can be used directly by compute nodes as directories, files, and mount points. Changes between the file system and S3 are synchronized automatically by the service in the background.&lt;br&gt;&lt;br&gt;
With this architecture, S3 Files does not migrate the entire dataset. Instead, it only places a portion of the current working set into the high‑performance layer on demand, while the “source of truth” for the data remains in S3.&lt;/p&gt;
&lt;h2&gt;
  
  
  How S3 Files works: mount, import, and sync
&lt;/h2&gt;

&lt;p&gt;For S3 Files, mounting is only the beginning. What really matters is the data path after mounting: how the scope is determined, what gets imported on first access, which requests go to the high‑performance layer, and how writes are synced back to S3. These mechanisms directly determine the performance boundaries and cost structure discussed later.&lt;br&gt;&lt;br&gt;
The figure below shows &lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-files.html#s3-files-how-it-works" rel="noopener noreferrer"&gt;S3 Files mount architecture&lt;/a&gt;:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe7l9wr6qpsa6kazmr78r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe7l9wr6qpsa6kazmr78r.png" alt=" " width="800" height="493"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Taking an EC2 instance mounting an existing S3 bucket as an example, the real questions are not about the mount command itself, but about what happens after mounting: how data is imported, accessed, and synced. Here are the key technical details.&lt;/p&gt;
&lt;h3&gt;
  
  
  Scope: import the entire bucket or a specific prefix
&lt;/h3&gt;

&lt;p&gt;In S3, files are organized using &lt;a href="https://docs.aws.amazon.com/AmazonS3/latest/userguide/using-prefixes.html" rel="noopener noreferrer"&gt;prefixes&lt;/a&gt;. You can think of them as similar to “subdirectories” or “sub-buckets” within a bucket, but they are not actual directories. S3 Files can mount an entire S3 bucket (i.e., &lt;code&gt;s3://my-bucket/&lt;/code&gt;) or restrict the scope to a specific prefix (i.e., only &lt;code&gt;s3://my-bucket/data/ml/&lt;/code&gt;). This is especially important for huge S3 buckets containing millions or even hundreds of millions of objects, because a broader scope adds more metadata sync overhead and potentially more operational complexity.&lt;br&gt;&lt;br&gt;
When using S3 Files on a compute node, AWS provides a custom mount client &lt;code&gt;amazon-efs-utils&lt;/code&gt;. When mounting, you use the file system ID assigned by AWS for the S3 Files instance, not the bucket name.&lt;br&gt;&lt;br&gt;
You can create a local mount directory and mount it using the dedicated &lt;code&gt;s3files&lt;/code&gt; file system type as shown below:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;yum &lt;span class="nt"&gt;-y&lt;/span&gt; &lt;span class="nb"&gt;install &lt;/span&gt;amazon-efs-utils
&lt;span class="nb"&gt;sudo mkdir&lt;/span&gt; /mnt/s3files
&lt;span class="nb"&gt;sudo &lt;/span&gt;mount &lt;span class="nt"&gt;-t&lt;/span&gt; s3files file-system-id:/ /mnt/s3files
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When mounting, you can specify a subdirectory (or actually a prefix) in the mount path (for example, &lt;code&gt;file-system-id:/data/ml/&lt;/code&gt;). However, this only affects the local mount point, as the underlying metadata sync still applies to the entire bucket, and other mount points could inadvertently access other subdirectories. Very often in practice, a better approach is to limit the scope to a specific prefix when creating the S3 Files instance, rather than using the entire bucket and trying to restrict access afterward.&lt;/p&gt;

&lt;h3&gt;
  
  
  First access: import triggers and size thresholds
&lt;/h3&gt;

&lt;p&gt;S3 Files does not immediately move the entire dataset into the high‑performance layer after mounting. Data import is triggered by access events. The default mode is &lt;code&gt;trigger=ON_DIRECTORY_FIRST_ACCESS&lt;/code&gt;: when you first access a directory, the system imports the metadata of files under that directory and asynchronously moves the data of small files that meet the criteria into the EFS high‑performance layer.  &lt;/p&gt;

&lt;p&gt;If configured as &lt;code&gt;trigger=ON_FILE_ACCESS&lt;/code&gt;, the first directory traversal only imports metadata; data enters the high‑performance layer only when a file is actually read for the first time. This saves space and import cost, but first‑read latency is higher.  &lt;/p&gt;

&lt;p&gt;The most critical control parameter here is &lt;code&gt;sizeLessThan&lt;/code&gt;, which takes a value in bytes. By default, only files smaller than 128 KiB (i.e., &lt;code&gt;sizeLessThan=131072&lt;/code&gt;) are moved into the high‑performance layer during import. Larger files only have their metadata imported; their content is still retrieved mainly from S3. In other words, S3 Files optimizes primarily for small files and low‑latency access, not for warming up all data into the high‑performance layer. For AI training datasets that consist mainly of 10 MB‑level images or video files, this is crucial: even after a directory traversal, those large files may not actually enter the high‑performance layer under default settings.&lt;/p&gt;

&lt;h3&gt;
  
  
  Sync intervals and conflict resolution
&lt;/h3&gt;

&lt;p&gt;S3 Files automatically maintains bidirectional sync between the file system and S3 in the background. When data in S3 changes, the file system view updates accordingly. Writes from compute nodes first land in the EFS high‑performance layer and then are batch‑synced back to S3 asynchronously. By default, the system aggregates modifications for a period before writing back.  &lt;/p&gt;

&lt;p&gt;The conflict resolution principle is clear: S3 is always the source of truth. If a file system modification hasn’t yet been synced back and the corresponding object has been updated by another application in S3, the system uses the latest version from S3 and moves the conflicting file to a &lt;code&gt;.s3files-lost+found-*&lt;/code&gt; directory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Performance boundaries and cost structure of S3 Files
&lt;/h2&gt;

&lt;p&gt;The previous section explained how S3 Files works. This section discusses the resulting performance boundaries and cost structure. High‑performance layer occupancy, large‑file read path, write flow, and the amplification effects of partial updates and directory operations are the four most important aspects to evaluate.&lt;/p&gt;

&lt;h3&gt;
  
  
  Occupancy, eviction, and cost of the EFS high‑performance layer
&lt;/h3&gt;

&lt;p&gt;S3 Files does not use capacity‑based LRU eviction. Instead, it uses access‑time‑based lifecycle management. By default, data that has been synced to S3 and not read for 30 days is evicted from the EFS high‑performance layer. This period is controlled by &lt;code&gt;daysAfterLastAccess&lt;/code&gt; (configurable from 1 to 365 days).  &lt;/p&gt;

&lt;p&gt;This means the cost depends on how much data needs to reside in EFS and for how long. If your working set is large and stays active for a long time, costs will rise accordingly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Large‑file direct reads and random reads: essentially client passthrough
&lt;/h3&gt;

&lt;p&gt;S3 Files does not route all large‑file reads through the EFS high‑performance layer. With the default &lt;code&gt;sizeLessThan&lt;/code&gt; of 128 KiB, only files below that threshold have their data moved into the high‑performance layer during import. For files already synced to S3, reads of 128 KiB or larger are streamed directly from S3.  &lt;/p&gt;

&lt;p&gt;In other words, S3 Files focuses on optimizing small‑file access and low latency, not on keeping large‑file reads in the high‑performance layer.  &lt;/p&gt;

&lt;p&gt;This direct‑read path requires that the compute resources themselves have permission to read the source bucket. AWS documentation explicitly requires roles to have &lt;code&gt;s3:GetObject&lt;/code&gt; and &lt;code&gt;s3:GetObjectVersion&lt;/code&gt; permissions; otherwise, the client cannot read directly from S3.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cost of sequential writes: large‑scale writes introduce additional flow costs
&lt;/h3&gt;

&lt;p&gt;The write path of S3 Files is not directly to S3. All writes first go into the EFS high‑performance layer and then are asynchronously synced back to S3.  &lt;/p&gt;

&lt;p&gt;This means that if your workload continuously generates large amounts of result data – for example, sequentially writing hundreds of terabytes of training outputs or analysis results – that data will incur two extra costs when flowing through S3 Files:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Data flow cost: Writes go into the high‑performance layer, then are later synced to S3. Compared to writing directly to S3, this path has an additional intermediate flow overhead.
&lt;/li&gt;
&lt;li&gt;Short‑term residency cost: After data is synced, it’s not immediately evicted from the high‑performance layer. It stays until the eviction condition is met. So temporary data from large writes may occupy EFS capacity for some time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Using the current AWS US East (N. Virginia) regional pricing as an example: writing to EFS is about $0.06/GB, and reading it back for sync to S3 is about $0.03/GB. For the data flow alone, each 1 TB written adds roughly $90 in extra cost. If that data continues to reside in EFS after sync, further high‑performance layer storage costs will apply.  &lt;/p&gt;

&lt;p&gt;This is why S3 Files is better suited for reading existing data than for sustaining large‑scale, continuous result writing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Partial updates and directory operations: amplification due to the object model
&lt;/h3&gt;

&lt;p&gt;S3 Files does not chunk data at the file system level. It tries to maintain a direct mapping between files and S3 objects. The drawback: when you perform a small random write or append to a large file – from the application’s perspective a tiny update – the backend sync back to S3 can easily amplify into a full object write and versioning overhead.  &lt;/p&gt;

&lt;p&gt;For example, a user appends a 100 KB image key to a 100 GB LMDB file through S3 Files. The application sees a small write. But this modification is not immediately written back to S3. It’s aggregated for up to  60 seconds and then synced to the bucket. Unlike block storage, which would only update a single block, this type of update is more likely to be amplified into an object write, sync latency, and version storage costs. The larger the file and the more frequent the modifications, the more serious this issue becomes.  &lt;/p&gt;

&lt;p&gt;Directory renaming is similarly constrained by S3’s flat namespace. S3 does not have native directory metadata. Therefore, when you perform &lt;code&gt;rename&lt;/code&gt; or &lt;code&gt;mv&lt;/code&gt;, S3 Files cannot just update a single metadata entry. Instead, it must write new objects for every file under the directory and delete the old ones. For a directory with tens of millions of objects, this significantly prolongs sync time and increases S3 request costs. Until the sync completes, the file system view and the S3 view may be temporarily inconsistent.  &lt;/p&gt;

&lt;p&gt;Overall, S3 Files’ strengths are native access, zero data migration, and good compatibility with existing S3 assets. The drawback, however, is that once the workload shifts to large-file reads, continuous writes, frequent partial updates, or operations on large directories, both performance and cost are amplified more quickly. For this reason, S3 Files is better suited for lightweight shared access scenarios, while in heavy workloads such as training, data production, and large-scale analytics, its drawbacks tend to emerge earlier.&lt;/p&gt;

&lt;h2&gt;
  
  
  JuiceFS vs. S3 Files: two different architectural philosophies
&lt;/h2&gt;

&lt;p&gt;As we saw earlier, many of S3 Files’ boundaries are not accidental; they are typical of solutions that try to map files directly to S3 objects. Whether it’s the early s3fs, Mountpoint for Amazon S3 (which focuses on high‑throughput reads), or today’s S3 Files, all of them try to maintain a direct mapping between files and S3 objects in exchange for transparent access to existing S3 data.  &lt;/p&gt;

&lt;p&gt;The advantage of this approach is transparency and low‑effort adoption. The cost has inherent limitations due to the S3 object model. That’s why directory operations tend to devolve into per‑object requests, and partial updates on large files easily amplify into write amplification, sync delays, and extra costs.  &lt;/p&gt;

&lt;p&gt;This is precisely why the difference between JuiceFS and these solutions is not a matter of a single feature or metric, but a fundamental divergence in architectural philosophy. JuiceFS is not an access layer that “mounts S3 as a file system”; it’s a cloud‑native distributed file system built on top of object storage. It adopts a decoupled metadata and data architecture: file data is stored in the underlying object storage, while metadata is independently managed by a high‑performance key‑value store. This makes JuiceFS much better suited for heavier production workloads such as training, analytics, and data production.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmz8vaakt4nte1dkav0ql.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmz8vaakt4nte1dkav0ql.png" alt=" " width="800" height="604"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;To help you compare these two architectures, here is a comprehensive table comparing JuiceFS and S3 Files:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Comparison basis&lt;/th&gt;
&lt;th&gt;JuiceFS&lt;/th&gt;
&lt;th&gt;Amazon S3 Files&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Architecture&lt;/td&gt;
&lt;td&gt;Metadata and data decoupled; files are chunked and written to object storage&lt;/td&gt;
&lt;td&gt;EFS‑based proxy; data is not chunked; 1:1 mapping based on file size (&amp;lt; 128 KB by default)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Core cost&lt;/td&gt;
&lt;td&gt;Open source software; costs from object storage, metadata engine, and cache resources&lt;/td&gt;
&lt;td&gt;In addition to S3 storage, pay for EFS storage ($ 0.30/GB) plus data sync fees&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Read/Write amplification (random writes)&lt;/td&gt;
&lt;td&gt;Very low in many cases. Because of chunking, a partial random write typically updates only the affected block, not the whole file&lt;/td&gt;
&lt;td&gt;Can be very high in data production scenarios. Random modification of a large file leads to rewriting and retransmitting the entire object&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tiering strategy&lt;/td&gt;
&lt;td&gt;Based on capacity and access heat; hot data is automatically warmed up into local disk/memory cache on compute nodes&lt;/td&gt;
&lt;td&gt;Based on file size and access time. Small files (&amp;lt; 128 KB) are cached in EFS; large files bypass EFS and read/write directly to S3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Small‑file performance&lt;/td&gt;
&lt;td&gt;Relies on a fully in‑memory metadata engine (like Redis or TiKV); well suited for many small files and metadata operations&lt;/td&gt;
&lt;td&gt;Depends on EFS performance and the NFS protocol&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Large‑file throughput&lt;/td&gt;
&lt;td&gt;Can use local NVMe / memory cache to boost throughput&lt;/td&gt;
&lt;td&gt;Depends on EFS gateway or S3 direct performance; large‑scale parallel throughput is tied to capacity quotas&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache consistency&lt;/td&gt;
&lt;td&gt;Strong consistency (close‑to‑open) enforced by the independent metadata service&lt;/td&gt;
&lt;td&gt;NFS close‑to‑open. But when concurrent modifications conflict between S3 and the file system, local EFS data is discarded into lost+found and S3 is forced as the source of truth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;POSIX compatibility&lt;/td&gt;
&lt;td&gt;Nearly 100% compatible. Supports hard links, atomic rename, and full lock semantics&lt;/td&gt;
&lt;td&gt;Subset of NFSv4.1/4.2. Does not support hard links or atomic rename&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Permission management&lt;/td&gt;
&lt;td&gt;Supports standard POSIX permissions, ACLs, extended ACLs, and more&lt;/td&gt;
&lt;td&gt;Supports standard POSIX permissions, ACLs, extended ACLs, and more&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Encryption and security&lt;/td&gt;
&lt;td&gt;In‑transit encryption and at‑rest encryption&lt;/td&gt;
&lt;td&gt;In‑transit TLS encryption and at‑rest SSE‑KMS encryption&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AI‑specific optimizations&lt;/td&gt;
&lt;td&gt;Deeply optimized for AI‑common formats (such as LMDB and Safetensors) with mmap reads and local data warm-up&lt;/td&gt;
&lt;td&gt;No AI‑specific optimizations; relies on basic streaming reads&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;There is no silver bullet. The right choice depends on your specific scenario.  &lt;/p&gt;

&lt;p&gt;The launch of S3 Files fills a gap in the AWS official ecosystem: seamless, zero‑migration native conversion of S3 into a file system. Its design philosophy is clear: 100% format transparency with the S3 object ecosystem, with targeted optimizations for small‑file (&amp;lt;128 KB) read/write performance in AI scenarios.&lt;/p&gt;

&lt;h3&gt;
  
  
  When to choose S3 Files
&lt;/h3&gt;

&lt;p&gt;If your core need is to allow legacy applications, shell scripts, or traditional software to access existing S3 data as files without changing your architecture, or if you need a general shared file space that is primarily used for read‑only, small‑file, sequential workloads, then S3 Files is the more natural choice. Its native managed service, plug‑and‑play, and zero data migration can significantly lower the barrier to entry (though you may need to pay high EFS storage and sync costs for that convenience).&lt;/p&gt;

&lt;h3&gt;
  
  
  When to choose JuiceFS
&lt;/h3&gt;

&lt;p&gt;If your workloads involve AI model training, data production, high‑performance computing (HPC), or big data analytics – facing tens of millions of small files, random reads/writes of terabyte‑scale large files, or demanding higher mmap performance, cache hit rates, and overall throughput – then JuiceFS is the better fit. Compared to S3 Files, JuiceFS’ data chunking, independent metadata engine, and more flexible caching system make it better suited for heavy‑workload, long‑running production and AI training/inference file system scenarios.  &lt;/p&gt;

&lt;p&gt;If you have any questions for this article, feel free to join &lt;a href="https://github.com/juicedata/juicefs/discussions/" rel="noopener noreferrer"&gt;JuiceFS discussions on GitHub&lt;/a&gt; and &lt;a href="http://go.juicefs.com/discord" rel="noopener noreferrer"&gt;community on Discord&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>JuiceFS 1.4: Lower Costs, Faster Metadata, and Better Control for Massive Data Management</title>
      <dc:creator>DASWU</dc:creator>
      <pubDate>Wed, 15 Jul 2026 09:04:24 +0000</pubDate>
      <link>https://dev.to/daswu/juicefs-14-lower-costs-faster-metadata-and-better-control-for-massive-data-management-1nno</link>
      <guid>https://dev.to/daswu/juicefs-14-lower-costs-faster-metadata-and-better-control-for-massive-data-management-1nno</guid>
      <description>&lt;p&gt;&lt;a href="https://github.com/juicedata/juicefs/releases/tag/v1.4.0" rel="noopener noreferrer"&gt;JuiceFS Community Edition 1.4&lt;/a&gt; is released. It’s the fifth major release since the open source edition was introduced in 2021 and is now the new Long-Term Support (LTS) version. We’ll continue to maintain both v1.4 and v1.3, while v1.2 has reached end of maintenance.&lt;/p&gt;

&lt;p&gt;JuiceFS has now surpassed &lt;strong&gt;14.2k &lt;a href="https://github.com/juicedata/juicefs" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; stars&lt;/strong&gt;. According to anonymous usage statistics reported by users, the total amount of data managed by JuiceFS Community Edition has exceeded &lt;strong&gt;1.4 EB&lt;/strong&gt;, representing more than &lt;strong&gt;700× growth&lt;/strong&gt; since 2022.&lt;/p&gt;

&lt;p&gt;As JuiceFS is increasingly adopted for large-scale data management, high-concurrency workloads, and multi-user shared environments, long-standing challenges such as storage cost optimization, metadata performance, and resource governance have become more prominent. These are the primary focus areas of the 1.4 release.&lt;/p&gt;

&lt;p&gt;In this post, we'll walk through the key improvements in JuiceFS 1.4, including tiered storage, faster metadata operations, enhanced resource management, more reliable data synchronization, metadata change tracking, and broader platform compatibility.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lower storage costs with file- and directory-level tiered storage
&lt;/h2&gt;

&lt;p&gt;As file systems continue to grow, different datasets naturally diverge in access frequency, performance requirements, and retention periods. Using a single storage type uniformly makes it difficult to simultaneously meet the performance needs of frequently accessed data and the cost-control requirements of infrequently accessed data. &lt;a href="https://en.wikipedia.org/wiki/Object_storage" rel="noopener noreferrer"&gt;Object storage&lt;/a&gt; typically offers different storage classes based on access patterns, including hot data, warm (infrequent access) data, and cold (archival) data storage.  &lt;/p&gt;

&lt;p&gt;JuiceFS has supported setting object storage types via &lt;code&gt;--storage-class&lt;/code&gt; since v1.1, but the configuration granularity was mainly at the file system default or mount point level. &lt;strong&gt;JuiceFS 1.4 integrates storage class into the file system semantics, supporting storage tier settings on a per-file or per-directory basis.&lt;/strong&gt; Directory-level configurations can be inherited by subsequently created files and subdirectories. This facilitates tiered management by project, dataset, or application directory.  &lt;/p&gt;

&lt;p&gt;The storage tiers can be configured flexibly according to the object storage vendor being used. When writing new data, JuiceFS writes it to the corresponding object storage type based on the configuration of the file or its parent directory. For existing data, you can also adjust the metadata configuration and leverage the data migration capabilities on the object storage side to move it to a new storage tier. This capability is suitable for scenarios such as AI training datasets, log archiving, backup data, historical experiment data, and offline analysis results. For archive storage, it’s still necessary to evaluate retrieval latency and fees. For more implementation details, usage methods, and future evolution, see &lt;a href="https://juicefs.com/en/blog/engineering/juicefs-tiered-storage" rel="noopener noreferrer"&gt;A Deep Dive into JuiceFS 1.4 Tiered Storage&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Faster metadata operations: batch delete, batch clone, and hotspot read optimization
&lt;/h2&gt;

&lt;p&gt;In workloads involving massive numbers of small files, large directories, or high-concurrency access, metadata operations often become the primary performance bottleneck.  &lt;/p&gt;

&lt;p&gt;JuiceFS 1.4 addresses write transaction overhead and hotspot read overhead in metadata operations with optimizations including batch delete, batch clone, and &lt;a href="https://redis.io/docs/latest/develop/clients/client-side-caching/" rel="noopener noreferrer"&gt;Redis client-side caching&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Batch delete and clone: reducing transaction overhead
&lt;/h3&gt;

&lt;p&gt;Previously, deleting a large number of files required the system to process them one by one, sequentially updating directory entries, inodes, space statistics, trash, and quota metadata. JuiceFS 1.4 consolidates the deletion of multiple non-directory files within the same directory into a batch transaction, reducing the repetitive overhead of per-file operations. This is applicable to scenarios like large directory cleanup, temporary data reclamation, training sample cleanup, and log directory deletion.&lt;/p&gt;

&lt;p&gt;Batch cloning targets directory replication and snapshot scenarios. &lt;a href="https://juicefs.com/docs/community/guide/clone/" rel="noopener noreferrer"&gt;&lt;code&gt;juicefs clone&lt;/code&gt;&lt;/a&gt; does not copy the underlying data blocks but creates new file records at the metadata layer and reuses the source file's data block references. JuiceFS 1.4 further reduces the metadata transactions generated by per-file cloning by processing clones of multiple files within the same directory in batch. This is ideal for AI dataset version management, experiment environment preparation, and large-scale directory snapshots.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F52nsn96kdqfq5r9kp34r.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F52nsn96kdqfq5r9kp34r.png" alt=" " width="800" height="572"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F84q9bjjcohu9hruvycly.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F84q9bjjcohu9hruvycly.png" alt=" " width="800" height="602"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Redis client-side caching: reducing hotspot metadata read overhead
&lt;/h3&gt;

&lt;p&gt;In high-concurrency reads, path resolution, directory entry lookups, and file attribute queries generate a large number of repeated requests. When Redis is used as the metadata engine, these requests require round trips between the client and Redis. This may impact access latency and increase Redis load.  &lt;/p&gt;

&lt;p&gt;&lt;strong&gt;JuiceFS 1.4 caches hot inode attributes and directory entries locally on the client side. When cached, it can reduce repeated queries to Redis.&lt;/strong&gt; When related metadata changes, the local state is updated through a cache invalidation mechanism. It's important to note that this capability caches metadata, not file content.  &lt;/p&gt;

&lt;p&gt;It’s particularly beneficial for read-heavy workloads with stable hot paths, such as AI training data loading, large-scale container startup, and multi-task concurrent reading. For more implementation details, see &lt;a href="https://juicefs.com/en/blog/engineering/improve-metadata-operation-performance" rel="noopener noreferrer"&gt;Faster Metadata Operations with Batch Unlink, Batch Clone, and Redis Client-Side Caching&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Improved resource management: user quotas and trash usage statistics
&lt;/h2&gt;

&lt;p&gt;In distributed storage environments, storage resources are often shared among multiple users, teams, and projects.&lt;/p&gt;

&lt;p&gt;Without effective governance, accidental writes or abnormal workloads from a single user can quickly consume large amounts of storage space or inodes. This affects both system stability and operating costs. Quota management is a critical means of establishing predictable resource boundaries in shared environments.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;JuiceFS Community Edition 1.4 introduces &lt;a href="https://juicefs.com/docs/community/guide/quota/#user-and-group-quota" rel="noopener noreferrer"&gt;user and group quotas&lt;/a&gt;, allowing administrators to monitor, configure, and enforce resource limits based on identities.&lt;/strong&gt;  &lt;/p&gt;

&lt;p&gt;Resource governance is now extended beyond file system and directory quotas to include user- and group-level quotas, making it especially suitable for shared clusters and AI training platforms.  &lt;/p&gt;

&lt;p&gt;To reduce metadata overhead in multi-client environments, JuiceFS uses asynchronous accounting so that usage statistics converge gradually over time. For details, see &lt;a href="https://juicefs.com/en/blog/engineering/quota-design-in-distributed-architecture" rel="noopener noreferrer"&gt;Quota Design in Distributed Architectures: Implementation and Use Cases in JuiceFS.&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Supported quota types include:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Quota type&lt;/th&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;th&gt;Design goal&lt;/th&gt;
&lt;th&gt;Typical use case&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Total file system quota&lt;/td&gt;
&lt;td&gt;Entire file system&lt;/td&gt;
&lt;td&gt;Prevents overall resource runaway&lt;/td&gt;
&lt;td&gt;Cost budget control, capacity limit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Subdirectory quota&lt;/td&gt;
&lt;td&gt;Directory subtree&lt;/td&gt;
&lt;td&gt;Blocks abnormal write behavior&lt;/td&gt;
&lt;td&gt;Prevents misoperations, small‑file storms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User quota&lt;/td&gt;
&lt;td&gt;Per user&lt;/td&gt;
&lt;td&gt;Isolates impact between different applications&lt;/td&gt;
&lt;td&gt;Multi‑tenant data management&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User group quota&lt;/td&gt;
&lt;td&gt;Project or department&lt;/td&gt;
&lt;td&gt;Cost allocation and team limits&lt;/td&gt;
&lt;td&gt;Shared environment for AI projects&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;JuiceFS 1.4 also improves trash space visibility.  &lt;/p&gt;

&lt;p&gt;Deleted files may remain in the trash for a retention period, making it difficult to understand why storage space has not yet been reclaimed. The enhanced &lt;code&gt;summary&lt;/code&gt; tool now reports trash usage, helping administrators identify storage consumption and make informed cleanup, retention, or expansion decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Expanded capabilities: sync, backup, and change tracking
&lt;/h2&gt;

&lt;h3&gt;
  
  
  More reliable large-scale sync
&lt;/h3&gt;

&lt;p&gt;Large-scale migration, cross-cloud sync, backup, and archival workloads often face interruptions, security requirements, and bandwidth contention.  &lt;/p&gt;

&lt;p&gt;JuiceFS 1.4 significantly enhances &lt;a href="https://juicefs.com/docs/community/guide/sync/" rel="noopener noreferrer"&gt;&lt;code&gt;juicefs sync&lt;/code&gt;&lt;/a&gt; with three major capabilities:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Resumable sync:&lt;/strong&gt; It reduces recovery costs after task interruptions. During synchronization, JuiceFS records the task progress. If the task exits abnormally or is manually interrupted, it can resume from the saved state, reducing repeated scanning and processing. This capability is suitable for migration and backup scenarios with a large number of objects, long task durations, or unstable cross-cloud links.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data encryption and decryption:&lt;/strong&gt; In cross-cloud backup and archiving scenarios, client-side encryption is a common compliance requirement. &lt;strong&gt;JuiceFS 1.4 supports completing encrypted writes, decryption recovery, or re-encryption within the sync pipeline&lt;/strong&gt;, reducing reliance on external encryption tools. This capability is suitable for off-site backup, sensitive data migration, key rotation, and compliance auditing. However, it requires careful management of key storage and recovery processes.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Global traffic control:&lt;/strong&gt; It provides bandwidth constraints for concurrent multiple sync tasks. Compared to per-process rate limiting, version 1.4 can centrally manage the overall bandwidth usage of multiple sync tasks, reducing the impact of sync tasks on online application and other network activities. This is suitable for cross-cloud transfers, multi-task concurrent backups, data center migrations, and shared outbound link scenarios. For implementation details, see &lt;a href="https://juicefs.com/en/blog/engineering/resumable-sync-encryption-bandwidth-control" rel="noopener noreferrer"&gt;JuiceFS Sync for PB-Scale Data Transfers: Resumable Sync, Encryption, and Bandwidth Control&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6il3ea8pfsnt0sv8lrx9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6il3ea8pfsnt0sv8lrx9.png" alt=" " width="800" height="361"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Changelog: metadata changes become traceable
&lt;/h3&gt;

&lt;p&gt;JuiceFS Community Edition 1.4 introduces a &lt;a href="https://juicefs.com/docs/community/administration/changelog/" rel="noopener noreferrer"&gt;metadata changelog&lt;/a&gt; capability, which records metadata change events across the file system.  &lt;/p&gt;

&lt;p&gt;Previously, troubleshooting relied primarily on client-side access logs, which only reflected operations performed through individual mount points. In multi-client deployments, reconstructing a complete sequence of events was often difficult.  &lt;/p&gt;

&lt;p&gt;A changelog records metadata operations—including file creation, deletion, attribute updates, and renames—directly at the metadata layer, providing a unified source for troubleshooting, auditing, and incremental processing.&lt;br&gt;&lt;br&gt;
Administrators can now quickly identify accidental deletions, unexpected renames, permission changes, and metadata modifications without collecting logs from every client.  &lt;/p&gt;

&lt;p&gt;When issues like accidental deletion, abnormal renaming, or unexpected permission or attribute changes occur, administrators can review the relevant change records based on the changelog. This reduces dependence on single-client logs and shortens the troubleshooting path. It also provides a more unified source of metadata changes for operational auditing.  &lt;/p&gt;

&lt;p&gt;In backup, migration, and recovery scenarios, the changelog can serve as a reference for incremental processing. For large-scale file systems, numerous changes may occur between two full backups or migration tasks. By recording the metadata changes during this period, the changelog can provide input for subsequent incremental backups, migrations, or recovery processes, reducing reliance on full scans.&lt;/p&gt;

&lt;h2&gt;
  
  
  Better support across different environments
&lt;/h2&gt;

&lt;p&gt;JuiceFS Community Edition 1.4 further improves compatibility across diverse deployment environments.  &lt;/p&gt;

&lt;p&gt;On Windows clients, the release improves cross-platform consistency and stability, including user mapping, permission mapping, and file access behavior, reducing compatibility issues when Linux and Windows clients access the same file system.  &lt;/p&gt;

&lt;p&gt;For the Java SDK and the Hadoop ecosystem, &lt;strong&gt;JuiceFS 1.4 adds Kerberos authentication, completing support for Hadoop secure mode&lt;/strong&gt;.  &lt;/p&gt;

&lt;p&gt;JuiceFS 1.3 already introduced Apache Ranger integration for authorization and access control. Together, Kerberos authenticates who the user is, while Ranger determines what the user can access, providing a more complete security model for enterprise big data platforms.  &lt;/p&gt;

&lt;p&gt;On the storage backend side, JuiceFS 1.4 also adds support for SMB/CIFS-based storage. This makes it easier to integrate with existing NAS or file-sharing infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Continued growth in scale and AI adoption
&lt;/h2&gt;

&lt;p&gt;According to anonymous usage statistics, JuiceFS Community Edition now powers nearly 70,000 file systems managing over 1.4 EB of data, with deployment scale continuing to grow.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwh534flxj6g30gpubim8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwh534flxj6g30gpubim8.png" alt=" " width="800" height="315"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Over the past year, AI applications have continued to expand from model training to inference services, agents, and multi-cloud scheduling, placing higher demands on data storage. These changes are also reflected in the use cases shared by community users, covering areas such as &lt;a href="https://en.wikipedia.org/wiki/Large_language_model" rel="noopener noreferrer"&gt;large language models&lt;/a&gt;, autonomous driving, quantitative investment, and computing platforms. We thank these users for sharing their real-world practices, providing valuable references for more teams building AI data infrastructure.&lt;br&gt;&lt;br&gt;
New user stories:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI training and large language models&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://juicefs.com/en/blog/user-stories/artificial-intelligence-model-training-unified-storage-solution" rel="noopener noreferrer"&gt;INTSIG Built Unified Storage Based on JuiceFS to Support Petabyte-Scale AI Training&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://juicefs.com/en/blog/user-stories/artificial-intelligence-storage-large-language-model-multimodal" rel="noopener noreferrer"&gt;StepFun Built an Efficient and Cost-Effective LLM Storage Platform with JuiceFS&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://juicefs.com/en/blog/user-stories/artificial-intelligence-big-data-cloud-native-storage" rel="noopener noreferrer"&gt;JuiceFS at Xiaomi: Unified Storage for AI, Big Data, and Cloud‑Native Workloads&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;AIGC&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://juicefs.com/en/blog/user-stories/multi-cloud-storage-artificial-intelligence-training" rel="noopener noreferrer"&gt;Why Gaoding Technology Chose JuiceFS for AI Storage in a Multi-Cloud Architecture&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://juicefs.com/en/blog/user-stories/aigc-storage-glusterfs-cephfs-vs-juicefs" rel="noopener noreferrer"&gt;From GlusterFS to JuiceFS: Lightillusions Achieved 2.5x Faster 3D AIGC Data Processing&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Autonomous driving and robotics&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://juicefs.com/en/blog/user-stories/multi-cloud-storage-autonomous-driving" rel="noopener noreferrer"&gt;Zelos Tech Manages Hundreds of Millions of Files for Autonomous Driving with JuiceFS&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://juicefs.com/en/blog/user-stories/multi-cloud-store-massive-small-files" rel="noopener noreferrer"&gt;How D-Robotics Manages Massive Small Files in a Multi-Cloud Environment with JuiceFS&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Inference and AI agents&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://juicefs.com/en/blog/user-stories/ai-storage-model-distribution-cross-cloud-inference" rel="noopener noreferrer"&gt;How Gongjiyun Keeps Model Distribution Fast Enough for Cross-Cloud Elastic Inference&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://juicefs.com/en/blog/user-stories/multi-cloud-storage-ai-agent" rel="noopener noreferrer"&gt;42x Faster Writes &amp;amp; 85% Throughput Gain: JuiceFS for Multi-Cloud AI Agents&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Quant investment&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://juicefs.com/en/blog/user-stories/quantitative-storage-artificial-intelligence-solution" rel="noopener noreferrer"&gt;JuiceFS+MinIO: Ariste AI Achieved 3x Faster I/O and Cut Storage Costs by 40%+&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Big data&lt;/strong&gt;  &lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://juicefs.com/en/blog/user-stories/juicefs-vs-alluxio-ai-storage-naver" rel="noopener noreferrer"&gt;NAVER, Korea's No.1 Search Engine, Chose JuiceFS over Alluxio for AI Storage&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Development of JuiceFS 1.4 spanned nearly a year. During the release cycle, the community reported &lt;strong&gt;366 issues&lt;/strong&gt;, merged &lt;strong&gt;515 pull requests&lt;/strong&gt;, and welcomed contributions from &lt;strong&gt;59 contributors&lt;/strong&gt;. We sincerely thank everyone who reported issues, contributed code, improved documentation, and helped evolve JuiceFS for increasingly demanding production environments. Your participation drives the rapid growth of JuiceFS.&lt;/p&gt;

&lt;p&gt;Download and try JuiceFS 1.4 &lt;a href="https://github.com/juicedata/juicefs/releases/tag/v1.4.0" rel="noopener noreferrer"&gt;here&lt;/a&gt;.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>JuiceFS Sync for PB-Scale Data Transfers: Resumable Sync, Encryption, and Bandwidth Control</title>
      <dc:creator>DASWU</dc:creator>
      <pubDate>Fri, 10 Jul 2026 06:56:21 +0000</pubDate>
      <link>https://dev.to/daswu/juicefs-sync-for-pb-scale-data-transfers-resumable-sync-encryption-and-bandwidth-control-1lda</link>
      <guid>https://dev.to/daswu/juicefs-sync-for-pb-scale-data-transfers-resumable-sync-encryption-and-bandwidth-control-1lda</guid>
      <description>&lt;p&gt;In scenarios such as data migration, cross-cloud synchronization, and object storage backup, &lt;a href="https://juicefs.com/docs/community/guide/sync/" rel="noopener noreferrer"&gt;&lt;code&gt;juicefs sync&lt;/code&gt;&lt;/a&gt; is commonly used to transfer large volumes of data. When datasets grow to the TB- or PB-scale, with millions or even billions of objects, a single synchronization task may run for hours or even days.&lt;/p&gt;

&lt;p&gt;As these long-running jobs progress, several common challenges tend to emerge:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;After network interruptions, process crashes, or node restarts, tasks often struggle to resume from a consistent state and may need to rescan or reprocess data.
&lt;/li&gt;
&lt;li&gt;Backup workflows may expose plaintext data and face compliance or security requirements.
&lt;/li&gt;
&lt;li&gt;When multiple sync jobs run concurrently, bandwidth contention becomes significant, while the overall transfer process lacks effective global control.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To address these challenges, &lt;a href="https://github.com/juicedata/juicefs/releases/tag/v1.4.0" rel="noopener noreferrer"&gt;JuiceFS 1.4&lt;/a&gt; introduces three major enhancements to &lt;code&gt;sync&lt;/code&gt;: resumable sync, data encryption/decryption, and global traffic control.&lt;/p&gt;

&lt;p&gt;In this article, we’ll explain the use cases, implementation details, and configuration methods for each feature.&lt;/p&gt;

&lt;h2&gt;
  
  
  Resumable sync
&lt;/h2&gt;

&lt;p&gt;In earlier versions, if a synchronization task failed or was interrupted, rerunning &lt;code&gt;juicefs sync&lt;/code&gt; required rescanning both the source and destination before determining which objects had already been synchronized and which still needed to be copied.&lt;/p&gt;

&lt;p&gt;For workloads involving hundreds of millions of objects or large files, the scan itself could incur substantial time and object-storage request costs.&lt;/p&gt;

&lt;p&gt;To address this issue, JuiceFS 1.4 introduces a &lt;a href="https://juicefs.com/docs/community/guide/sync/#checkpoint" rel="noopener noreferrer"&gt;&lt;strong&gt;resumable sync&lt;/strong&gt;&lt;/a&gt; mechanism for &lt;code&gt;sync&lt;/code&gt;. When enabled, synchronization progress is periodically saved to the destination. If the task is interrupted, rerunning the same command automatically locates and loads the matching checkpoint and resumes from the last unfinished position, avoiding a full restart.&lt;/p&gt;

&lt;h3&gt;
  
  
  How it works
&lt;/h3&gt;

&lt;p&gt;When resumable sync is enabled, &lt;code&gt;sync&lt;/code&gt; stores a JSON state file on the destination side:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;.juicefs-sync-checkpoint.&amp;lt;&lt;span class="nb"&gt;hash&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;&amp;lt;hash&amp;gt;&lt;/code&gt; value is derived from the source, destination, and key synchronization parameters. This ensures that a task only loads checkpoints created for itself, preventing accidental reuse across different jobs.  &lt;/p&gt;

&lt;p&gt;The workflow is shown below:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhgjy6bwcjgy4vr48tyw5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhgjy6bwcjgy4vr48tyw5.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Checkpoint save, restore, and cleanup workflow in &lt;code&gt;juicefs sync&lt;/code&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;When &lt;code&gt;sync&lt;/code&gt; starts, it first looks for a checkpoint matching the current task.
&lt;/li&gt;
&lt;li&gt;If a matching checkpoint is found, execution resumes from the saved state. Otherwise, synchronization starts normally with a fresh scan. &lt;code&gt;sync&lt;/code&gt; traverses multiple prefixes concurrently, maintaining independent state for each prefix, including:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- Whether traversal is complete  
- The last scanned position  
- Pending objects to synchronize  
- Failed objects  
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;When restoring from a checkpoint:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;- Pending and failed objects recorded in the checkpoint are re-added to the task queue.  
- Prefixes that were not fully traversed resume scanning from their saved positions.  
- Fully traversed prefixes only continue processing unfinished objects recorded in the checkpoint.  
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;During execution, progress is saved asynchronously at a configurable interval, which defaults to every &lt;strong&gt;10 seconds&lt;/strong&gt;.  &lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;After successful completion, the checkpoint file is automatically removed. If the task fails or is interrupted, the checkpoint is retained for resumption on the next execution of the same command.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvqe3izhnrkzyxy4sm1e4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvqe3izhnrkzyxy4sm1e4.png" alt=" " width="799" height="454"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In cluster mode, only a single checkpoint exists and is maintained centrally by Manager.&lt;br&gt;&lt;br&gt;
Workers do not directly read or write checkpoint files on the destination. Instead, they:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Pull tasks from Manager
&lt;/li&gt;
&lt;li&gt;Execute synchronization
&lt;/li&gt;
&lt;li&gt;Report results back to Manager&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Manager aggregates completed objects, failed objects, statistics, and multipart-upload state into the global checkpoint.&lt;/p&gt;
&lt;h3&gt;
  
  
  Usage
&lt;/h3&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Enable resumable sync.&lt;/span&gt;
juicefs &lt;span class="nb"&gt;sync&lt;/span&gt; &lt;span class="nt"&gt;--enable-checkpoint&lt;/span&gt; SRC DST

&lt;span class="c"&gt;# Customize checkpoint save interval (default: 10s).&lt;/span&gt;
juicefs &lt;span class="nb"&gt;sync&lt;/span&gt; &lt;span class="nt"&gt;--enable-checkpoint&lt;/span&gt; &lt;span class="nt"&gt;--checkpoint-interval&lt;/span&gt; 30s SRC DST

&lt;span class="c"&gt;# Ignore existing checkpoints and restart from scratch.&lt;/span&gt;
juicefs &lt;span class="nb"&gt;sync&lt;/span&gt; &lt;span class="nt"&gt;--enable-checkpoint&lt;/span&gt; &lt;span class="nt"&gt;--checkpoint-force-reset&lt;/span&gt; SRC DST
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;h2&gt;
  
  
  Data encryption and decryption
&lt;/h2&gt;

&lt;p&gt;For cross-cloud backup and archival workflows, client-side encryption is often required to satisfy compliance requirements such as data sovereignty, encryption at rest, and secure migration of sensitive data.&lt;/p&gt;

&lt;p&gt;Previously, &lt;code&gt;juicefs sync&lt;/code&gt; did not provide built-in encryption capabilities. Users who wanted to write encrypted data to the destination typically had to use external tools for additional processing.&lt;/p&gt;

&lt;p&gt;In JuiceFS 1.4, &lt;a href="https://juicefs.com/docs/community/guide/sync/#encryption-and-decryption" rel="noopener noreferrer"&gt;streaming encryption and decryption&lt;/a&gt; are integrated directly into the synchronization pipeline, enabling three common workflows:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Encrypt-on-write:&lt;/strong&gt; Encrypt plaintext data before writing it to the destination, suitable for encrypted backup and archiving.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decrypt-on-read:&lt;/strong&gt; Read encrypted data from the source and write decrypted data to the destination, suitable for data recovery or plaintext migration.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-encryption:&lt;/strong&gt; Decrypts source data with an old key and re-encrypts it with a new key before writing to the destination, suitable for key rotation or cryptographic algorithm migration.&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;
  
  
  Chunk-based streaming encryption
&lt;/h3&gt;

&lt;p&gt;To support object storage Range GET operations while avoiding excessive memory usage for large files all at once, &lt;code&gt;sync&lt;/code&gt; uses a fixed-size 1 MiB chunk-based streaming encryption scheme.&lt;/p&gt;

&lt;p&gt;A file is first divided into plaintext chunks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="o"&gt;[&lt;/span&gt;chunk 1: 1 MiB][chunk 2: 1 MiB] ... &lt;span class="o"&gt;[&lt;/span&gt;chunk N: ≤1 MiB]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each plaintext chunk is encrypted independently.&lt;/p&gt;

&lt;p&gt;Each encrypted chunk consists of a 4-byte header and the ciphertext data, where the 4-byte header stores the actual ciphertext length (&lt;code&gt;ct_len&lt;/code&gt;):&lt;/p&gt;

&lt;p&gt;Each encrypted block: [4B ct_len][ciphertext + padding]&lt;/p&gt;

&lt;p&gt;Encrypted file: [encrypted chunk 1][encrypted chunk 2] ... [encrypted chunk N]&lt;/p&gt;

&lt;p&gt;The encrypted block size is determined by the plaintext chunk size plus encryption overhead: &lt;code&gt;plainChunkSize + overhead&lt;/code&gt;. The &lt;code&gt;plainChunkSize&lt;/code&gt; is fixed at 1 MiB, and the &lt;code&gt;overhead&lt;/code&gt; depends on the encryption algorithm and key type used.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1o43z6voxk33foonm2ak.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1o43z6voxk33foonm2ak.png" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This design allows random reads to retrieve only the required encrypted chunk rather than downloading the entire file. Because encrypted objects contain additional headers, padding, and encryption metadata, the destination object is typically larger than the original plaintext file.&lt;/p&gt;

&lt;h3&gt;
  
  
  Supported algorithms
&lt;/h3&gt;

&lt;p&gt;The table below shows the supported algorithms:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;Symmetric cipher&lt;/th&gt;
&lt;th&gt;Key encapsulation&lt;/th&gt;
&lt;th&gt;Typical use case&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;aes256gcm-rsa (default)&lt;/td&gt;
&lt;td&gt;AES-256-GCM&lt;/td&gt;
&lt;td&gt;RSA&lt;/td&gt;
&lt;td&gt;General-purpose workloads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;chacha20-rsa&lt;/td&gt;
&lt;td&gt;ChaCha20-Poly1305&lt;/td&gt;
&lt;td&gt;RSA&lt;/td&gt;
&lt;td&gt;Environments without efficient AES hardware acceleration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;sm4gcm&lt;/td&gt;
&lt;td&gt;SM4-GCM&lt;/td&gt;
&lt;td&gt;SM2&lt;/td&gt;
&lt;td&gt;Scenarios requiring Chinese commercial cryptography standards&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Usage
&lt;/h3&gt;

&lt;p&gt;The following examples use RSA keys.&lt;/p&gt;

&lt;p&gt;Generate a key pair:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Generate an RSA private key (the public key is derived automatically).&lt;/span&gt;
openssl genrsa &lt;span class="nt"&gt;-out&lt;/span&gt; private.pem 2048

&lt;span class="c"&gt;# Generate a password-protected private key.&lt;/span&gt;
openssl genrsa &lt;span class="nt"&gt;-aes256&lt;/span&gt; &lt;span class="nt"&gt;-out&lt;/span&gt; private.pem 2048
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Scenario 1: Encrypt and write to destination&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;juicefs &lt;span class="nb"&gt;sync&lt;/span&gt; /local/data s3://mybucket/backup 
    &lt;span class="nt"&gt;--encrypt-rsa-key&lt;/span&gt; /path/to/private.pem
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Scenario 2: Decrypt and read from source for data recovery or plaintext migration&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;juicefs &lt;span class="nb"&gt;sync &lt;/span&gt;s3://mybucket/backup /local/data 
    &lt;span class="nt"&gt;--decrypt-rsa-key&lt;/span&gt; /path/to/private.pem
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Scenario 3: Re-encrypt for key rotation or algorithm migration&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Decrypt data encrypted with the old key and re-encrypt with the new key to new storage.&lt;/span&gt;
juicefs &lt;span class="nb"&gt;sync &lt;/span&gt;s3://old-bucket/encrypted s3://new-bucket/re-encrypted 
    &lt;span class="nt"&gt;--decrypt-rsa-key&lt;/span&gt; /path/to/old-private.pem 
    &lt;span class="nt"&gt;--encrypt-rsa-key&lt;/span&gt; /path/to/new-private.pem
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the private key is password-protected, the password can be provided via environment variables:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# For encryption scenarios, use JFS_ENCRYPT_RSA_PASSPHRASE.&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;JFS_ENCRYPT_RSA_PASSPHRASE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-passphrase"&lt;/span&gt;
juicefs &lt;span class="nb"&gt;sync&lt;/span&gt; /local/data s3://mybucket/backup &lt;span class="nt"&gt;--encrypt-rsa-key&lt;/span&gt; private.pem

&lt;span class="c"&gt;# For decryption scenarios, use JFS_DECRYPT_RSA_PASSPHRASE.&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;JFS_DECRYPT_RSA_PASSPHRASE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-passphrase"&lt;/span&gt;
juicefs &lt;span class="nb"&gt;sync &lt;/span&gt;s3://mybucket/backup /local/data &lt;span class="nt"&gt;--decrypt-rsa-key&lt;/span&gt; private.pem
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Encrypted data is stored using a JuiceFS-specific format and can only be decrypted through &lt;code&gt;juicefs sync&lt;/code&gt; with the corresponding key.
&lt;/li&gt;
&lt;li&gt;Back up encryption keys carefully. Once a private key is lost, encrypted data cannot be recovered.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Global traffic control
&lt;/h2&gt;

&lt;p&gt;In earlier versions, &lt;code&gt;juicefs sync&lt;/code&gt; already supported per-process rate limiting via &lt;a href="https://juicefs.com/docs/community/guide/sync/#global-traffic-control" rel="noopener noreferrer"&gt;&lt;code&gt;--bwlimit&lt;/code&gt;&lt;/a&gt;. However, when multiple sync processes run concurrently—such as multiple Workers in a distributed sync, or multiple independent sync tasks sharing the same egress link—per-process limiting cannot constrain total bandwidth usage. The egress link may still be saturated, affecting other application traffic.&lt;/p&gt;

&lt;p&gt;JuiceFS 1.4 introduces the &lt;code&gt;--traffic-control-url&lt;/code&gt; parameter. Multiple sync processes can connect to the same external traffic control service, which allocates bandwidth quotas uniformly, enabling cross-process, cross-task global rate limiting.&lt;/p&gt;

&lt;h3&gt;
  
  
  How it works
&lt;/h3&gt;

&lt;p&gt;Global traffic control uses a &lt;strong&gt;token bucket model&lt;/strong&gt;. Before transmitting data, each sync process requests byte credits from the same traffic-control service.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F96m4p1e9fucs9tt262zo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F96m4p1e9fucs9tt262zo.png" alt=" " width="800" height="525"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Each process periodically requests a certain number of bytes (credit) before data transfer.&lt;/p&gt;

&lt;p&gt;The traffic-control service determines:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;How many bytes to grant
&lt;/li&gt;
&lt;li&gt;How long the granted quota remains valid&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When credits are exhausted, the process requests additional credits.&lt;/p&gt;

&lt;p&gt;If a quota is about to expire before being fully consumed, the unused portion is returned to the service in advance.&lt;/p&gt;

&lt;p&gt;The service exposes a simple HTTP API for granting and reclaiming quotas. This must be implemented by the user or integrated with an existing service:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;POST /traffic-control
Content-Type: application/json

Request:
&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"bytes"&lt;/span&gt;: 1048576&lt;span class="o"&gt;}&lt;/span&gt;
  bytes &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; 0: Request byte credits.
  bytes &amp;lt; 0: Return unused credits.


Response:
&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;"granted"&lt;/span&gt;: 524288, &lt;span class="s2"&gt;"expired"&lt;/span&gt;: 1000&lt;span class="o"&gt;}&lt;/span&gt;
  granted: Number of bytes granted this time.
  expired: Credit validity period &lt;span class="o"&gt;(&lt;/span&gt;milliseconds&lt;span class="o"&gt;)&lt;/span&gt;&lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;During synchronization, &lt;code&gt;sync&lt;/code&gt; requests quotas from the traffic control service before transmitting data. If no credits are available, transmission blocks until new credits are obtained. In this way, multiple sync tasks can share a single global bandwidth limit, preventing the total traffic from becoming uncontrolled even when individual tasks have their own limits.&lt;/p&gt;

&lt;h3&gt;
  
  
  Usage
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Deploy a traffic-control service first.&lt;/span&gt;
&lt;span class="c"&gt;# (Example: listen on port 8080 and cap total bandwidth at 100 Mbps)&lt;/span&gt;
&lt;span class="c"&gt;# (Service implementation is user-defined; JuiceFS only calls the API)&lt;/span&gt;

&lt;span class="c"&gt;# Multiple sync processes share the same control service.&lt;/span&gt;
juicefs &lt;span class="nb"&gt;sync &lt;/span&gt;SRC1 DST1 &lt;span class="nt"&gt;--traffic-control-url&lt;/span&gt; http://127.0.0.1:8080/traffic-control &amp;amp;
juicefs &lt;span class="nb"&gt;sync &lt;/span&gt;SRC2 DST2 &lt;span class="nt"&gt;--traffic-control-url&lt;/span&gt; http://127.0.0.1:8080/traffic-control &amp;amp;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--traffic-control-url&lt;/code&gt; can be combined with &lt;code&gt;--bwlimit&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The two mechanisms are independent:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;--bwlimit&lt;/code&gt; limits the bandwidth of a single sync process.
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;--traffic-control-url&lt;/code&gt; limits aggregate bandwidth across multiple processes.
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Per-process limit: 50 Mbps. All processes combined respect the service-side cap.&lt;/span&gt;
juicefs &lt;span class="nb"&gt;sync &lt;/span&gt;SRC DST 
    &lt;span class="nt"&gt;--bwlimit&lt;/span&gt; 50 
    &lt;span class="nt"&gt;--traffic-control-url&lt;/span&gt; http://controller:8080/traffic-control
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;JuiceFS 1.4 enhancements to &lt;code&gt;sync&lt;/code&gt; include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Resumable sync&lt;/strong&gt; reduces recovery costs after task interruptions.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Encryption and decryption&lt;/strong&gt; improve the security of backups and archival data.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Global traffic control&lt;/strong&gt; enables multiple synchronization tasks to share bandwidth in a coordinated manner.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For scenarios such as data migration, cross-cloud sync, object storage backup, and encrypted archiving, users can combine these capabilities flexibly based on task scale, network environment, and security requirements.&lt;/p&gt;

&lt;p&gt;If you have any questions for this article, feel free to join &lt;a href="https://github.com/juicedata/juicefs/discussions/" rel="noopener noreferrer"&gt;JuiceFS discussions on GitHub&lt;/a&gt; and &lt;a href="http://go.juicefs.com/discord" rel="noopener noreferrer"&gt;community on Discord&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Monitoring JuiceFS with Better Stack</title>
      <dc:creator>DASWU</dc:creator>
      <pubDate>Fri, 26 Jun 2026 07:46:14 +0000</pubDate>
      <link>https://dev.to/daswu/monitoring-juicefs-with-better-stack-18ce</link>
      <guid>https://dev.to/daswu/monitoring-juicefs-with-better-stack-18ce</guid>
      <description>&lt;p&gt;After deployment, &lt;a href="https://juicefs.com/docs/community/introduction/" rel="noopener noreferrer"&gt;JuiceFS&lt;/a&gt; feels like a local drive, but underneath it's a sophisticated distributed system. This perfectly reflects one of its core design principles: distributed systems are complex, but from a user's perspective, they should be simple to use.&lt;/p&gt;

&lt;p&gt;Even so, that simplicity on the surface doesn't negate the need for deep visibility. For any critical storage system, gaining real-time visibility into its operations is crucial to prevent subtle performance degradations from escalating into significant incidents.&lt;/p&gt;

&lt;p&gt;Fortunately, JuiceFS exposes a suite of monitoring metrics, including throughput, IOPS, latency, data size, and many more, in the widely adopted Prometheus format, making it ready for modern monitoring stacks. Traditionally, you would probably pair Prometheus with Grafana to collect these metrics and visualize them. This is indeed a powerful combination. However, deploying, managing, and maintaining these systems yourself adds operational overhead again. Ironically, you may want to monitor them too, and trust me, you would rather not create yet another monitoring stack just to monitor your Prometheus and Grafana combo.&lt;/p&gt;

&lt;p&gt;That's where &lt;a href="https://betterstack.com/" rel="noopener noreferrer"&gt;Better Stack&lt;/a&gt; comes in. It is a fully managed SaaS observability platform that combines user-friendly dashboards, tracing, logging, error tracking, incident management, automatic alerting, and even &lt;strong&gt;AI-powered SRE&lt;/strong&gt;, all for a predictable, cost-effective price. With Better Stack, you get the power of the best-in-class tools out of the box without the operational overhead.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frffw4cxuurkbuwep0szt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frffw4cxuurkbuwep0szt.png" alt=" " width="799" height="569"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In this post, we'll guide you through setting up a comprehensive monitoring system for JuiceFS using Better Stack, from metric ingestion to intelligent alerting, so you can ensure your file system remains healthy and performant.&lt;/p&gt;

&lt;h2&gt;
  
  
  Preparing the JuiceFS file system
&lt;/h2&gt;

&lt;p&gt;Before diving into setting up Better Stack for monitoring, you'll need an existing JuiceFS file system that is actively publishing metrics. &lt;a href="https://juicefs.com/docs/community/introduction/" rel="noopener noreferrer"&gt;JuiceFS Community Edition&lt;/a&gt; and &lt;a href="https://juicefs.com/docs/cloud/" rel="noopener noreferrer"&gt;JuiceFS Enterprise Edition&lt;/a&gt; (our cloud service is based on JuiceFS Enterprise Edition) both expose real-time status metrics in Prometheus format, but they do it in slightly different ways.&lt;/p&gt;

&lt;p&gt;For the JuiceFS Community Edition, after mounting the file system, JuiceFS automatically exposes metrics via &lt;code&gt;http://localhost:9567/metrics&lt;/code&gt; by default on the mounting host where the JuiceFS client is running. You can customize this port using the &lt;code&gt;--metrics&lt;/code&gt; option if needed.&lt;/p&gt;

&lt;p&gt;On the other hand, for JuiceFS Enterprise Edition &amp;amp; Cloud Service, metrics are exposed through the console via dedicated API endpoints. You'll need to replace &lt;code&gt;VOLUME_NAME&lt;/code&gt; with your file system name and &lt;code&gt;API_TOKEN&lt;/code&gt; with your API token. In this case, both Prometheus and JSON formats are available for metrics:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Prometheus: &lt;code&gt;https://juicefs.com/api/vol/VOLUME_NAME/metrics?token=API_TOKEN&lt;/code&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;JSON: &lt;code&gt;https://juicefs.com/api/volume/VOLUME_NAME/status?token=YOUR_TOKEN&lt;/code&gt;&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A quick but important note: &lt;strong&gt;metrics are only generated when the file system is mounted&lt;/strong&gt;. So before proceeding, ensure your JuiceFS file system is properly mounted and accessible. In this guide, we will use the JuiceFS Cloud Service, as it's the simplest to get started. If you haven't set up JuiceFS yet, please refer to the &lt;a href="https://juicefs.com/docs/cloud/getting_started" rel="noopener noreferrer"&gt;documentation&lt;/a&gt; for detailed instructions. Once you have created the first file system, URLs for the metrics mentioned above would be available under its &lt;strong&gt;Monitor&lt;/strong&gt; tab.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo71ci0b6hafbe2bcnc5h.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo71ci0b6hafbe2bcnc5h.png" alt=" " width="800" height="404"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Setting up a metrics source in Better Stack
&lt;/h2&gt;

&lt;p&gt;With your JuiceFS file system up and running (don't forget to mount the file system to a host machine) and publishing metrics, the next step is to configure Better Stack to start ingesting that data.&lt;/p&gt;

&lt;p&gt;First, if you haven't already, register for a Better Stack account. The process is seamless. Using a work email is recommended, and the platform provides clear guidance to help you set up your account and organization.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6zmxybq3ao4kgx5o4ekp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6zmxybq3ao4kgx5o4ekp.png" alt=" " width="800" height="405"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once you're logged in, follow these steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;In the left-hand navigation panel, head to &lt;strong&gt;Telemetry&lt;/strong&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Under the &lt;strong&gt;Sources&lt;/strong&gt; section, click &lt;strong&gt;Connect source&lt;/strong&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Give your telemetry data source a descriptive name, such as "jfs-better-stack" or "juicefs-production", to easily identify it later.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fezbvtp6qflgzoaw29wk6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fezbvtp6qflgzoaw29wk6.png" alt=" " width="800" height="405"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Now, you'll configure how Better Stack should collect your metrics. In the collector settings:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Under &lt;strong&gt;Metrics&lt;/strong&gt;, choose the &lt;strong&gt;Prometheus scrape&lt;/strong&gt; option and click &lt;strong&gt;Connect source&lt;/strong&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;In the &lt;strong&gt;URLs to scrape&lt;/strong&gt; section, input the JuiceFS metrics endpoint as described above.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgqs367692qv9qv9ntm1m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgqs367692qv9qv9ntm1m.png" alt=" " width="800" height="405"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Note that if you are not using the JuiceFS Cloud Service and your JuiceFS endpoint is behind a firewall, you'll need to allow traffic from Better Stack's scrape servers. The list of IP addresses to add to the allowlist is available in their documentation and from &lt;a href="https://telemetry.betterstack.com/prometheus-scrape-ips.txt" rel="noopener noreferrer"&gt;here&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;After saving the configuration, Better Stack will begin scraping the endpoint. Your JuiceFS metrics should be received within a few seconds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Creating a dashboard with AI SRE
&lt;/h2&gt;

&lt;p&gt;With your JuiceFS metrics flowing into Better Stack, it's time to visualize them. You could build a dashboard manually, but Better Stack provides a smarter and more efficient way to do it by using &lt;strong&gt;AI SRE&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is AI SRE?
&lt;/h3&gt;

&lt;p&gt;AI SRE (Site Reliability Engineering) is Better Stack's chat-based site reliability assistant. It's an autonomous AI agent that can read your telemetry data, analyze incidents, build dashboards, and even write code to fix errors. Instead of waiting for humans to manually set up charts and queries, AI SRE can generate comprehensive dashboards for you based on a prompt.&lt;/p&gt;

&lt;p&gt;It's notable that AI SRE is a paid feature. If you're on the free plan, you can still create dashboards manually using the drag-and-drop chart builder.&lt;/p&gt;

&lt;h3&gt;
  
  
  Creating a JuiceFS monitoring dashboard with a single prompt
&lt;/h3&gt;

&lt;p&gt;Once your metrics source is ready, follow these steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;From the left panel, head to &lt;strong&gt;Telemetry&lt;/strong&gt; and then &lt;strong&gt;Metrics&lt;/strong&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Click &lt;strong&gt;Create dashboard&lt;/strong&gt; and select the &lt;strong&gt;Create with AI&lt;/strong&gt; option.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;In the prompt field, give AI SRE a clear description of what you need. For example: "Create me a dashboard to track ALL JuiceFS metrics, such as latency, data size, etc."&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Also make sure to select the metrics &lt;strong&gt;Source&lt;/strong&gt; you created earlier (for example, "jfs-better-stack") so that AI SRE has the proper context and data to work with.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzuivi0hjlk7hskc8kke6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzuivi0hjlk7hskc8kke6.png" alt=" " width="800" height="403"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Give the platform a few minutes for the dashboard to be created. AI SRE will analyze your JuiceFS metrics and automatically generate a complete set of charts and panels for the important performance indicators such as throughput, IOPS, latency, and storage utilization. For my first time trying this, it just worked like a charm as shown below.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9l3she7awbnbozjdnhab.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9l3she7awbnbozjdnhab.png" alt=" " width="800" height="405"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;AI SRE is a powerful feature that does so much more than create dashboards. It can analyze incidents, perform root cause analysis, suggest fixes, and even open pull requests. We've only scratched the surface in this post. This is your first step toward a smarter, AI-assisted observability workflow. After building your dashboard, you can further customize it by adding panels, editing queries, or setting alerts directly from the graphs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;In this post, we have walked through how to build a complete observability system for JuiceFS with Better Stack. We started by setting up the JuiceFS file system and getting its Prometheus-formatted metrics, then created a metrics source in Better Stack to ingest the data. We examined rapid creation of a full dashboard with AI SRE.&lt;/p&gt;

&lt;p&gt;We hope this guide helps you gain better visibility into your JuiceFS deployment. If you have any questions or run into issues, we'd love to hear from you. Join the JuiceFS community on &lt;a href="https://github.com/juicedata/juicefs" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt; or &lt;a href="http://go.juicefs.com/discord" rel="noopener noreferrer"&gt;Discord&lt;/a&gt;. And don't forget to check out Better Stack's &lt;a href="https://betterstack.com/docs/getting-started/welcome/" rel="noopener noreferrer"&gt;documentation&lt;/a&gt; and their &lt;a href="https://www.youtube.com/@betterstack" rel="noopener noreferrer"&gt;amazing YouTube channel&lt;/a&gt; for practical insights about distributed file storage, observability, AI, and more.&lt;/p&gt;

</description>
      <category>opensource</category>
    </item>
    <item>
      <title>JuiceFS 1.4: Faster Metadata Operations with Batch Unlink, Batch Clone, and Redis Client-Side Caching</title>
      <dc:creator>DASWU</dc:creator>
      <pubDate>Thu, 18 Jun 2026 08:30:56 +0000</pubDate>
      <link>https://dev.to/daswu/juicefs-14-faster-metadata-operations-with-batch-unlink-batch-clone-and-redis-client-side-4ao1</link>
      <guid>https://dev.to/daswu/juicefs-14-faster-metadata-operations-with-batch-unlink-batch-clone-and-redis-client-side-4ao1</guid>
      <description>&lt;p&gt;In large-scale file access scenarios such as AI training and dataset management, metadata often becomes the first performance bottleneck as file counts and concurrency grow. Whether you're deleting millions of small files, cloning large datasets, or traversing directories under heavy concurrency, metadata performance directly impacts application efficiency.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://juicefs.com/docs/community/introduction" rel="noopener noreferrer"&gt;JuiceFS Community Edition&lt;/a&gt; 1.4 introduces three major metadata optimizations:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Batch unlink&lt;/strong&gt; for large-scale file deletion
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Batch clone&lt;/strong&gt; for metadata cloning
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Redis client-side caching&lt;/strong&gt; for hot metadata reads&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These improvements reduce transaction commits, network round trips, and redundant metadata lookups. In tests on a flat directory containing 100,000 files, batch unlink improved performance by up to &lt;strong&gt;93×&lt;/strong&gt;, while batch clone achieved up to &lt;strong&gt;24×&lt;/strong&gt; speedup.&lt;/p&gt;

&lt;p&gt;In this article, we’ll explain the motivation, design, and performance benefits behind these optimizations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deletion: From one‑by‑one to batched transactions
&lt;/h2&gt;

&lt;p&gt;Under &lt;a href="https://juicefs.com/docs/community/architecture" rel="noopener noreferrer"&gt;JuiceFS' metadata-data separation architecture&lt;/a&gt;, deleting a file involves much more than removing a directory entry. The system must also:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Update inode reference counts
&lt;/li&gt;
&lt;li&gt;Reclaim inode and space resources
&lt;/li&gt;
&lt;li&gt;Process trash entries
&lt;/li&gt;
&lt;li&gt;Update quota statistics&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These operations must typically be completed within the same transaction.&lt;/p&gt;

&lt;p&gt;When a directory contains hundreds of thousands or even millions of files, the traditional file-by-file deletion approach used by &lt;code&gt;rm -rf&lt;/code&gt; quickly becomes a bottleneck. Each &lt;code&gt;unlink&lt;/code&gt; request goes through the &lt;a href="https://www.kernel.org/doc/html/next/filesystems/fuse.html" rel="noopener noreferrer"&gt;FUSE protocol&lt;/a&gt;, switches between kernel and user space, and triggers a separate metadata transaction.&lt;/p&gt;

&lt;p&gt;As the number of files grows, the overhead from system calls, context switches, network round trips, and transaction commits accumulates rapidly.&lt;/p&gt;

&lt;p&gt;To mitigate this issue, JuiceFS previously introduced the &lt;code&gt;juicefs rmr&lt;/code&gt; command. Unlike &lt;code&gt;rm -rf&lt;/code&gt;, &lt;code&gt;rmr&lt;/code&gt; bypasses the FUSE layer and sends deletion requests directly to the client. It also supports multi-threaded deletion (50 threads by default), significantly improving throughput.&lt;/p&gt;

&lt;p&gt;However, each file deletion still requires its own metadata transaction. Deleting 100,000 files still means executing 100,000 transactions.&lt;/p&gt;

&lt;p&gt;Batch unlink takes optimization one step further by merging many independent deletion operations within the same directory into a single batch transaction, further removing network overhead.&lt;/p&gt;

&lt;h3&gt;
  
  
  Core design
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;The key is to turn many small transactions into fewer large ones. JuiceFS adds a batch unlink interface at the metadata engine layer. It allows the client to delete multiple non‑directory files under the same directory in one call.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When recursively clearing a directory, JuiceFS reduces deletion overhead in two ways:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Different subdirectories are handled concurrently with multi‑threaded deletion.
&lt;/li&gt;
&lt;li&gt;Inside each directory, normal files and symlinks are grouped into batches and sent to &lt;code&gt;BatchUnlink&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This merges many unlink operations into fewer batch transactions at the metadata level.&lt;br&gt;&lt;br&gt;
It's important to note that &lt;code&gt;BatchUnlink&lt;/code&gt; does not directly delete directories. Directory removal still follows the standard recursive workflow: empty the subdirectory first, and then delete the subdirectory itself.  Therefore, &lt;code&gt;BatchUnlink&lt;/code&gt; only applies to regular files and symbolic links within the same directory.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;This restriction preserves correct recursive deletion semantics while avoiding consistency risks to the directory tree structure.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn363fmtctinhgqmybcpp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn363fmtctinhgqmybcpp.png" alt=" " width="800" height="613"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;
  
  
  Implementation across metadata engines
&lt;/h3&gt;

&lt;p&gt;JuiceFS uses different batching strategies depending on the &lt;a href="https://juicefs.com/docs/community/databases_for_metadata/" rel="noopener noreferrer"&gt;metadata backend&lt;/a&gt; to minimize transaction commits and network round trips.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;SQL backends (MySQL, PostgreSQL, etc.):&lt;/strong&gt; Previously, each file deletion required its own sequence of &lt;code&gt;INSERT&lt;/code&gt;, &lt;code&gt;DELETE&lt;/code&gt;, and &lt;code&gt;UPDATE&lt;/code&gt; statements. With &lt;code&gt;BatchUnlink&lt;/code&gt;, the system:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Fetches all edge records for the target entries in a single batch query.
&lt;/li&gt;
&lt;li&gt;Retrieves the relevant inode attributes in a single locked batch query.
&lt;/li&gt;
&lt;li&gt;Executes edge deletions, inode state updates (decrementing nlink or marking for cleanup), and delfile entry insertions — all within one transaction.
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Instead of executing one transaction per file, the entire batch can now be completed in a single transaction.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Redis backend:&lt;/strong&gt; &lt;strong&gt;The optimization uses Redis pipelines and transactions.&lt;/strong&gt; Where individual deletions previously required separate command round trips, &lt;code&gt;BatchUnlink&lt;/code&gt; collects all &lt;code&gt;HDEL&lt;/code&gt; (dentry removal), &lt;code&gt;ZADD&lt;/code&gt; (enqueue for cleanup), &lt;code&gt;SET&lt;/code&gt; (inode attribute update), and &lt;code&gt;INCRBY&lt;/code&gt; (counter update) commands for multiple files into a single pipeline, executed atomically within one &lt;code&gt;MULTI&lt;/code&gt;/&lt;code&gt;EXEC&lt;/code&gt; transaction. To avoid blocking Redis' single-threaded event loop for too long, batch size is capped at 250 entries.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TiKV backend:&lt;/strong&gt; &lt;code&gt;BatchUnlink&lt;/code&gt; consolidates multiple deletions into a single transaction, using TiKV's batch write capability to reduce network round trips and transaction overhead. &lt;strong&gt;For distributed key-value backends, this kind of batching allows the backend's concurrent write capacity to be more fully utilized.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The figure below shows benchmark results on a flat directory of 100,000 files using &lt;code&gt;juicefs rmr --threads 16&lt;/code&gt;. &lt;code&gt;BatchUnlink&lt;/code&gt; delivers meaningful improvements across all metadata backends, with TiKV and Redis showing the largest gains.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqh8iyz1hsdtugcauahpt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqh8iyz1hsdtugcauahpt.png" alt=" " width="800" height="572"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Clone: From one‑by‑one copy to batched references
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://juicefs.com/docs/community/guide/clone/" rel="noopener noreferrer"&gt;&lt;code&gt;juicefs clone&lt;/code&gt;&lt;/a&gt; creates fast copies of files or directories for training dataset version management, experiment snapshots, and large-scale directory duplication. Its efficiency comes from the fact that cloning doesn't immediately copy the underlying data blocks. Instead, it creates new file records at the metadata layer and reuses the source file's existing block references. New data blocks are only allocated when the clone is actually written to. This avoids the time and storage overhead of a full copy.&lt;/p&gt;

&lt;p&gt;For large directory clones, the same problem as deletion arises: processing files one by one generates a large number of short transactions and network round trips. &lt;strong&gt;The core idea behind batch clone is to merge the clone operations for multiple files in the same directory into a single batch transaction.&lt;/strong&gt; When recursively cloning a directory, the system reads directory entries in batches as a stream. For each batch, all non-directory entries are collected and cloned together in one operation.&lt;/p&gt;

&lt;p&gt;One key implementation detail is &lt;strong&gt;inode pre-allocation&lt;/strong&gt;: before entering the transaction, the system uses &lt;code&gt;nextInode&lt;/code&gt; to pre-allocate target inodes for all entries to be cloned. This avoids lock contention from repeatedly requesting inodes inside the transaction. Once inside the transaction, the system batch-queries all source file attributes (with row locks), builds all the insertion data for target nodes, edges, chunks, symlinks, and xattrs, and then inserts everything in a single batch.&lt;/p&gt;

&lt;p&gt;Batch clone uses each backend's native batch write capabilities in a similar way to batch unlink. The per-backend implementation details won't be repeated here.&lt;/p&gt;

&lt;p&gt;The performance gains vary across backends depending on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Transaction models
&lt;/li&gt;
&lt;li&gt;Network communication overhead
&lt;/li&gt;
&lt;li&gt;Batch insertion efficiency for metadata records such as nodes, edges, and chunk references&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Results on a flat directory of 100,000 files are shown below. MySQL sees the largest improvement at approximately 24x; Redis at approximately 5x; TiKV at approximately 2x.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3ivuei7r7p7f27ej5trt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3ivuei7r7p7f27ej5trt.png" alt=" " width="800" height="602"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  Redis client-side caching: Keeping hot metadata local
&lt;/h2&gt;

&lt;p&gt;In high-concurrency metadata workloads such as AI training dataset access and large-scale container startup, network round trips between JuiceFS clients and Redis often become a major performance bottleneck.&lt;/p&gt;

&lt;p&gt;Consider the following operation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;open&lt;span class="o"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"/mnt/jfs/dataset/images/cat.jpg"&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Before the file can be opened, the Linux Virtual File System (VFS) must resolve every component in the path:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Look up &lt;code&gt;dataset&lt;/code&gt;.
&lt;/li&gt;
&lt;li&gt;Look up &lt;code&gt;images&lt;/code&gt;.
&lt;/li&gt;
&lt;li&gt;Look up &lt;code&gt;cat.jpg&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwscnpxkx0h3mq7lbuhfb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwscnpxkx0h3mq7lbuhfb.png" alt=" " width="799" height="445"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If the &lt;code&gt;images&lt;/code&gt; directory contains hundreds of thousands of files and training jobs perform random access across the dataset, each lookup requires a &lt;code&gt;GET&lt;/code&gt; request to Redis.&lt;br&gt;&lt;br&gt;
Under heavy concurrency, this results in large numbers of network round trips and increased Redis CPU utilization. &lt;strong&gt;Even though a single Redis query takes only a few dozen microseconds, network latency pushes each lookup to hundreds of microseconds or even milliseconds. When thousands of training processes are accessing files simultaneously, this overhead becomes significant.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  How it works: Redis 6.0 client-side caching
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://redis.io/docs/latest/develop/reference/client-side-caching/" rel="noopener noreferrer"&gt;Redis 6.0 introduced &lt;strong&gt;client-side caching&lt;/strong&gt;&lt;/a&gt;, which allows clients to cache hot keys locally and receive invalidation notifications whenever those keys are modified.&lt;/p&gt;

&lt;p&gt;Based on this capability, JuiceFS caches two categories of metadata in client memory:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Inode attribute cache.&lt;/strong&gt; Keyed by inode number, this stores the complete attribute data for a file, such as type, size, permissions, and timestamps. The caching is implemented transparently through hook mechanisms in the Redis driver layer. On query, it first checks the local cache; on hit, it returns immediately without any network request. On modification, it automatically invalidates the corresponding cache. Application logic requires no awareness of the cache.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Directory entry cache.&lt;/strong&gt; Keyed by "parent inode + path separator + filename," this caches the results of directory lookups. Unlike the inode attribute cache, the lookup logic for entry cache is embedded directly in the directory lookup path rather than being intercepted transparently at the driver layer. When entries for a directory are invalidated, all related cache entries under that directory are cleared using prefix matching. This allows path resolution and repeated access to hot entries in the same directory to be served from local memory.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Introducing client-side caching creates a consistency challenge in multi-mount scenarios.&lt;/strong&gt; When multiple clients share the same JuiceFS file system, an operation on one client — creating, deleting, renaming, or updating attributes of a file or directory — can invalidate cached inode attributes or directory entries on other clients. Without an effective invalidation mechanism, subsequent reads could hit stale metadata, causing the directory entries or file attributes seen by one client to diverge from the actual state in the backend.&lt;/p&gt;

&lt;p&gt;To address this, JuiceFS introduces a &lt;a href="https://redis.io/docs/latest/commands/client-tracking/" rel="noopener noreferrer"&gt;&lt;strong&gt;Tracking and Broadcast Invalidation&lt;/strong&gt;&lt;/a&gt; &lt;strong&gt;(BCAST)&lt;/strong&gt; model on top of Redis' client-side caching mechanism. After connecting to Redis, each client declares the metadata key prefixes it wants to track. When those keys are modified, Redis sends invalidation notifications to the relevant clients. On receiving a notification, the client clears the corresponding inode attribute cache or entry cache entries, so that subsequent accesses fetch fresh data from the metadata engine.&lt;/p&gt;

&lt;p&gt;In addition, at client initialization, JuiceFS warms up metadata for the root directory of the mount point. Since these files are typically the most frequently accessed, benchmarks show this warm-up significantly improves overall access performance.&lt;/p&gt;

&lt;p&gt;Through this mechanism, hot metadata can be reused locally. When the metadata changes, the related caches are evicted in time, reducing the risk of stale metadata.&lt;/p&gt;

&lt;h3&gt;
  
  
  When to use it
&lt;/h3&gt;

&lt;p&gt;Redis client‑side caching works best in read‑heavy, write‑light scenarios with repeated access to hot metadata. AI training dataset loading is a good example: the dataset is usually read‑only during training, and tasks repeatedly access the same directories and files, so inode attribute cache and entry cache hit often, reducing redundant lookups and remote metadata queries.&lt;/p&gt;

&lt;p&gt;The benefit is even more obvious when there is higher network latency between the client and the Redis metadata engine, such as in cross-availability-zone deployments.&lt;/p&gt;

&lt;p&gt;Redis 6.0 or later is required to use this feature. The default cache expiration time is 1 minute, which provides a safety net in case of network interruptions or connection anomalies where invalidation notifications may not arrive, preventing stale entries from persisting indefinitely. For workloads with stricter consistency requirements, the expiration time can be shortened or client-side caching can be disabled entirely to reduce the risk of reading stale metadata.&lt;/p&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;These three optimizations each target a different path through the metadata layer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Batch unlink&lt;/strong&gt; merges multiple independent unlink operations within the same directory into a single batch transaction.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Batch clone&lt;/strong&gt; merges multiple independent clone operations within the same directory into a single batch transaction.
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Redis client-side caching&lt;/strong&gt; keeps hot metadata in client memory, bringing read latency from network-level down to memory-level, with broadcast invalidation to maintain consistency across multiple clients.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;code&gt;BatchUnlink&lt;/code&gt; and &lt;code&gt;BatchClone&lt;/code&gt; are internal interfaces. Users do not call them directly. Just use the right commands: &lt;code&gt;juicefs rmr&lt;/code&gt; for deleting large directories, &lt;code&gt;juicefs clone&lt;/code&gt; for copying directories. The optimization will be applied automatically.&lt;/p&gt;

&lt;p&gt;One thing worth noting: both batch operations work by merging regular files within the same directory into a single batch transaction. Subdirectories are handled recursively by concurrent goroutines. The larger the directory, the greater the benefit.&lt;/p&gt;

&lt;p&gt;Batch operations mainly merge ordinary files under the same directory into one batch transaction. Subdirectories are handled recursively by concurrent goroutines. The larger the directory, the bigger the benefit.  &lt;/p&gt;

&lt;p&gt;All optimizations above are available in JuiceFS Community Edition 1.4. Upgrade the client to get the performance gains.  &lt;/p&gt;

&lt;p&gt;If you have any questions for this article, feel free to join &lt;a href="https://github.com/juicedata/juicefs/discussions/" rel="noopener noreferrer"&gt;JuiceFS discussions on GitHub&lt;/a&gt; and &lt;a href="http://go.juicefs.com/discord" rel="noopener noreferrer"&gt;community on Discord&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
    <item>
      <title>How Gongjiyun Keeps Model Distribution Fast Enough for Cross-Cloud Elastic Inference</title>
      <dc:creator>DASWU</dc:creator>
      <pubDate>Fri, 12 Jun 2026 03:25:32 +0000</pubDate>
      <link>https://dev.to/daswu/how-gongjiyun-keeps-model-distribution-fast-enough-for-cross-cloud-elastic-inference-2g8l</link>
      <guid>https://dev.to/daswu/how-gongjiyun-keeps-model-distribution-fast-enough-for-cross-cloud-elastic-inference-2g8l</guid>
      <description>&lt;p&gt;Founded in 2023 at Tsinghua University, &lt;a href="https://www.techinasia.com/companies/gongjiyun" rel="noopener noreferrer"&gt;Gongjiyun&lt;/a&gt; provides compute platforms and Model as a Service (MaaS) for artificial intelligence generated content (AIGC) enterprises and research institutions. We aim to alleviate the mismatch between elastic compute demand and supply. By aggregating idle IDC resources and edge resources, the platform offers containerized services, delivering rapidly schedulable compute for volatile workloads such as AI inference, video rendering, data processing, and data synthesis.&lt;/p&gt;

&lt;p&gt;In cross-cloud elastic inference scenarios, compute tasks can be scheduled to different regions, cloud environments, and clusters, but model files and application data are large and cannot be migrated as quickly as compute resources. Especially in online inference, the model repository is read‑heavy and frequently accessed – storage access performance directly affects service startup, elastic scaling, and request latency.&lt;/p&gt;

&lt;p&gt;To address this, we built an &lt;strong&gt;object storage acceleration&lt;/strong&gt; solution on top of &lt;a href="https://juicefs.com/docs/community/introduction/" rel="noopener noreferrer"&gt;JuiceFS&lt;/a&gt;, integrating users’ existing object storage into elastic inference clusters. Through a unified namespace, metadata import, FUSE mount, distributed cache, and data warm-up, it improves access efficiency for model repositories across clouds and clusters. In a case study with a leading text‑to‑image model community, the solution supports a tens‑of‑TB model repository, dynamic loading of checkpoints and low-rank adaptations (LoRAs), and elastic scaling of hundreds of GPUs at peak, while keeping additional latency within the customer’s acceptance range.&lt;/p&gt;

&lt;p&gt;In this post, we'll walk through why storage — not compute — is the real bottleneck in cross-cloud elastic inference, how we evaluated and chose JuiceFS, and the step-by-step optimizations that brought latency from +10s down to under 2s in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Elastic demand is widespread, but supply is hard to match
&lt;/h2&gt;

&lt;p&gt;As AI applications grow rapidly, compute demand continues to increase, but resource usage patterns differ across scenarios. &lt;strong&gt;Compared to training, which has stable resource needs, &lt;a href="https://www.ibm.com/think/topics/ai-inference" rel="noopener noreferrer"&gt;AI inference&lt;/a&gt;, data processing, and data synthesis are often more volatile&lt;/strong&gt;: office applications may see higher traffic during the day, entertainment apps during evenings or weekends, and project‑based data processing may consume large amounts of compute in short bursts then idle. For small teams or exploratory applications, elastic compute also helps them better evaluate the relationship between per‑request cost and application value.&lt;/p&gt;

&lt;p&gt;On the supply side, compute infrastructure is capital‑intensive. Resource providers are not incapable of offering elastic services, but they prefer long‑term dedicated leases to recover costs and reduce risk. As a result, low price, stability, and elasticity are difficult to achieve together:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dedicated leases are low‑cost and stable but lack elasticity.&lt;/li&gt;
&lt;li&gt;Spot resources are cheap and elastic but uncertain.&lt;/li&gt;
&lt;li&gt;On‑demand resources are elastic and stable but expensive.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In China, this contradiction is further reflected by a market dominated by dedicated leases, with elastic supply accounting for a small share.&lt;/p&gt;

&lt;p&gt;We aim to resolve this mismatch between elastic demand and supply. &lt;strong&gt;By aggregating idle IDC and edge resources, the platform offers containerized services, providing rapidly schedulable compute for AI inference, video rendering, data processing, and data synthesis.&lt;/strong&gt; At lower resource costs, we help users quickly spin up tasks during peaks, schedule them across clusters, and handle elastic demand, while enabling resource providers to improve utilization and monetize idle capacity beyond dedicated leases.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compute can be scheduled: How does storage keep up?
&lt;/h2&gt;

&lt;p&gt;As elastic compute platforms evolve, compute resource scheduling is easy. Container images can be synchronized across clusters via registries and distribution networks, tasks can be launched in different resource pools by schedulers, and traffic can be distributed via unified ingress and traffic management.&lt;/p&gt;

&lt;p&gt;But model and data files are typically large, making cross‑cloud, cross‑cluster migration costly and slow, unable to match the sub‑second startup and release of compute. Therefore, &lt;strong&gt;in cross‑cloud elastic inference architectures, the real limitation on system elasticity is often not compute scheduling, but the efficiency of data and model distribution&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Different application scenarios have different storage requirements:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;a href="https://www.ibm.com/think/topics/model-training" rel="noopener noreferrer"&gt;&lt;strong&gt;Model training&lt;/strong&gt;&lt;/a&gt;&lt;strong&gt;, development, and debugging:&lt;/strong&gt; These involve complex read‑write needs, including code repositories, model files, experiment results, and intermediate state. They also require high environment stability; users cannot tolerate state loss from frequent host switching. Thus, the platform typically provides long‑term stable compute resources and runtime environments, and storage needs can be met by existing stable storage systems.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Data processing:&lt;/strong&gt; This can be split further. If a single processing job has high application value and can cover cross‑cloud network transfer costs, you can build a pipeline that continuously pulls data from S3 or other object storage, processes it in the compute cluster, and writes back streaming. The system does not need large local storage. If the data scale is larger or per‑job value is low, local storage acts as a one‑time cache. Data flows through and does not need to be persisted.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What is truly more challenging is the online inference scenario&lt;/strong&gt;. Online inference services cannot tolerate downtime. However, the resources used by an elastic computing platform may come from idle resource pools. These resources could be preempted. Once resources in a certain data center or cluster become unavailable, the platform must be able to migrate tasks to other providers or other clusters in time. This means not only computing tasks must be migrated. Model files and related storage access capabilities must also be migrated at the same time&lt;/p&gt;

&lt;p&gt;Online inference has higher requirements for service continuity and cross-cluster migration capabilities, but its storage access pattern is also more clear. Compared to training, development, and debugging scenarios, inference workloads are typically read heavy. The core needs focus on efficient model loading, reading model weights, and accessing the model repository. For large models and online applications, model loading speed directly affects service startup time, elastic scaling efficiency, and request response stability. Therefore, inference scenarios are not suitable for simply adopting traditional read-write hybrid storage architectures. Instead, they are better suited for specialized optimizations around model distribution, read only access, and cache acceleration.&lt;/p&gt;

&lt;p&gt;In addition, an elastic computing platform usually does not host a user's complete application system. The user's primary cloud account, application database, model management system, and even some fixed computing resources often already exist in other clouds or on premises. For the platform to integrate with the user's application, it must be compatible with the user's existing model repository and model management processes. It cannot require the user to fully migrate the entire system.&lt;/p&gt;

&lt;p&gt;Therefore, &lt;strong&gt;to support cross-cloud elastic inference, we need more than just compute scheduling capabilities. We need a cross-cloud high-performance storage and model distribution solution tailored for model inference scenarios&lt;/strong&gt;. This solution must support hosting a large model repository and high-performance reading, it must adapt to the user's existing model management system. And it must provide stable data access capabilities when resources are migrated across clouds and clusters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why JuiceFS: Unified cross-cloud access, strongly consistent metadata, and high-performance cache
&lt;/h2&gt;

&lt;p&gt;Facing cross-cloud elastic inference scenarios, the storage system needs to meet several conditions at the same time:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;It must provide a unified access point across different clouds and clusters. It must support shared read-write access and unified metadata management.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;It must be compatible with the user's existing &lt;a href="https://en.wikipedia.org/wiki/Object_storage" rel="noopener noreferrer"&gt;object storage&lt;/a&gt; and model repository to avoid data migration.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;It needs low operational complexity and good read performance.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When evaluating storage options, we considered Ceph:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Ceph is mature. It’s suitable for building unified storage within a single data center or a stable resource domain.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;However, in cross cloud elastic inference scenarios, Ceph requires high network stability and operational skills. The overall integration cost is higher. So we did not choose it.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;We also evaluated Alluxio. However, in a &lt;a href="https://en.wikipedia.org/wiki/Multicloud" rel="noopener noreferrer"&gt;multi-cloud&lt;/a&gt; environment, multiple clusters need to access the same underlying object storage data concurrently. The workload is not purely read only; there are also occasional writes. This scenario requires strong data consistency. Therefore, Alluxio was not chosen for production.&lt;/p&gt;

&lt;p&gt;We finally chose JuiceFS mainly because:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;It uses object storage as the database.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;It provides a unified namespace and consistent file system view through an independent metadata service. This allows multiple clusters to access the same model data as a file system.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;This architecture is suitable for cross-cloud and cross-cluster model distribution and shared reading.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;It’s also compatible with the user's existing object storage and model repository, reducing data migration and application integration costs.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The decision to further adopt &lt;a href="https://juicefs.com/docs/cloud/" rel="noopener noreferrer"&gt;JuiceFS Enterprise Edition&lt;/a&gt; was mainly due to its &lt;strong&gt;distributed caching capabilities and managed metadata service&lt;/strong&gt;. In this scenario, the value of JuiceFS is not just providing a file system interface. It combines object storage, unified namespace, metadata management, and cache acceleration into a storage access layer that is better suited for cross-cloud elastic inference.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fn4bd67wtjwxd2zdb124p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fn4bd67wtjwxd2zdb124p.png" alt=" " width="800" height="594"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical: Object storage acceleration based on JuiceFS
&lt;/h2&gt;

&lt;p&gt;Based on JuiceFS, the platform encapsulates an object storage acceleration product. This product connects the user's existing object storage to the elastic inference cluster. It provides the storage as a high-performance file system for the application. The overall process is as follows.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Create a file system.&lt;/strong&gt; The user provides object storage access credentials, for example, AK/SK for S3-compatible storage. The credential permissions can be configured as read only or read-write based on application needs. The platform creates a corresponding JuiceFS file system based on that object storage.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Import metadata.&lt;/strong&gt; The platform uses the JuiceFS import feature to scan the metadata of files in object storage. Then, it imports that metadata into the JuiceFS metadata service. In this way, the model files originally stored by the user in object storage can be accessed as file system directories in JuiceFS.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Create a cache group.&lt;/strong&gt; Within each cluster that may host workloads, the platform sets up a JuiceFS cache group. This forms a distributed cache group. Before running a task, the platform can warm-up model files. It caches hot data in the target cluster in advance. This reduces the time needed to pull data from remote object storage when the inference service starts.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Mount to application Pods.&lt;/strong&gt; When the user's application runs, the platform uses the FUSE client to mount the JuiceFS file system into the application Pod. For the application, model files appear as local file system paths. Therefore, the original model reading logic usually does not need modification.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Enable node local cache.&lt;/strong&gt; Besides the cluster level cache group, the node where the FUSE client runs can also provide local cache. This improves repeated read and model loading performance. It further reduces direct access to remote object storage.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This object storage acceleration product essentially productizes the JuiceFS metadata import, distributed cache, data warm-up, and FUSE mounting process. It allows the user's existing object storage to serve cross-cloud inference tasks in a way that feels closer to a local file system.&lt;/p&gt;

&lt;p&gt;In addition, the JuiceFS cache group is independent from the file system access point. This characteristic, on one hand, adds management complexity on the platform side, because the platform needs to manage the relationships among the file system, cache groups, mount points, and task scheduling. On the other hand, it provides a foundation for cache isolation, independent scheduling, and fine-grained management based on clusters, users, or application scenarios in the future.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production case study: A leading text-to-image model community
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Scenario, challenges, and acceptance criteria
&lt;/h3&gt;

&lt;p&gt;One of the most representative cases in this object storage acceleration solution involves a leading Chinese text-to-image model community hosting tens of terabytes of model data, including large checkpoint base models and a larger number of smaller LoRA models. In practice, inference jobs typically load a checkpoint first, then load one or more LoRA models to perform combined inference.&lt;/p&gt;

&lt;p&gt;The company already operated compute infrastructure at scale — several thousand GPUs — but its workload, serving creative design and production use cases, exhibited significant variability. &lt;strong&gt;Overall average utilization was below 50%, yet during morning and afternoon peak hours on weekdays, load could reach 140% of normal capacity, degrading the user experience&lt;/strong&gt;. The customer therefore needed a highly elastic compute supply.&lt;/p&gt;

&lt;p&gt;We provided a high-elasticity resource model: compute support at the scale of hundreds of GPUs was available only during weekday peak hours — 10:00–12:00 AM and 2:00–6:00 PM — with resources scaling to zero at all other times.&lt;/p&gt;

&lt;p&gt;This meant the platform needed to provision hundreds of GPUs within a window of minutes, while consuming zero resources outside peak hours. For the customer, this model delivers large-scale compute during peak periods while avoiding payment for idle capacity. For the platform, it enables more efficient utilization and monetization of idle compute resources.&lt;/p&gt;

&lt;p&gt;The technical challenges were significant:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;A model repository of this scale cannot simply be replicated to every elastic cluster.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Inference services do not load all models once at startup. Model reads and switches happen continuously as user requests arrive, resulting in high access frequency. Therefore, the object storage acceleration solution needed to support not just large-scale model repository access, but stable read performance under continuous dynamic loading.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The customer's performance requirements were also strict. During acceptance testing, a portion of production traffic was routed to the elastic cluster. The requirement was that both the median and mean inference latency of the elastic cluster must stay within 2 seconds of the customer's own cluster. Given that individual inference jobs take on the order of tens of seconds, this requirement left virtually no room for additional latency introduced by the storage layer. In the first few rounds of testing, both median and mean inference latency on the elastic cluster exceeded the customer's own cluster by approximately 10 seconds — failing the acceptance criteria.&lt;/p&gt;

&lt;h3&gt;
  
  
  Performance optimization: Reducing additional latency on the elastic cluster
&lt;/h3&gt;

&lt;p&gt;Optimization began with the median. &lt;strong&gt;A high median indicates that a significant proportion of requests are experiencing performance degradation, not just a small number of outliers inflating the tail.&lt;/strong&gt; JuiceFS monitoring revealed that the cluster's cache hit rate was not reaching the expected level. In the current architecture, a cache miss requires a round trip over the public internet to the customer's object storage on Alibaba Cloud. This significantly increases model loading time and then affects inference request latency.&lt;/p&gt;

&lt;p&gt;To solve this, the platform used the isolation capability of the JuiceFS cache group. It assigned dedicated cache nodes to this customer, reserved enough cache space, and warmed up the core model data. After warming up, the access path for core models achieved nearly 100% cache hit rate. This effectively avoided the performance loss from cross public network backfilling.&lt;/p&gt;

&lt;p&gt;The second factor affecting the median was metadata access latency. Because the platform uses a unified cross-cluster architecture, the metadata service is accessed over the public internet, for example, via JuiceFS Cloud Service or a deployment on a remote host, and this latency affects overall model read performance.&lt;/p&gt;

&lt;p&gt;The platform took two measures to address this issue:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Enabling JuiceFS' open cache to keep metadata in local memory as much as possible.&lt;/strong&gt; Since this workload is predominantly read-only, caching is an effective way to reduce metadata access overhead.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Tuning the cluster's network rate-limiting policy&lt;/strong&gt;. While the platform cannot directly control network equipment in edge data centers, it can apply node-level rate limiting to prevent any single node from saturating the available bandwidth, improving overall network stability. After these optimizations, cluster-wide performance improved meaningfully and the median metric gradually reached the customer's requirement.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Once the median met the target, the mean still showed a gap. This indicated that long-tail requests remained, with a small number of requests taking significantly longer than normal and pulling up the overall average.&lt;/strong&gt; Further analysis traced this to node-level local cache — specifically, the FUSE cache quota. With limited cache capacity, the elastic cluster experienced more frequent cache evictions than the customer's own cluster, causing some requests to reload model data from scratch and increasing mean inference latency. The platform addressed this by increasing the FUSE local cache quota in the production environment, reducing eviction frequency, improving tail latency, and ultimately bringing the mean metric within acceptance. The system passed validation and has been running stably since.&lt;/p&gt;

&lt;h3&gt;
  
  
  Multi-tenant cache management
&lt;/h3&gt;

&lt;p&gt;After the single-tenant case was validated, the solution entered multi-tenant operation. As different tenants began time-sharing the same elastic nodes, a new issue emerged: cache contention between tenants.&lt;/p&gt;

&lt;p&gt;In the elastic resource model, FUSE clients do not actively clear node cache on exit. This is a reasonable design in single-tenant scenarios, where cached data from previous jobs can be reused by subsequent jobs to improve hit rates. &lt;strong&gt;However, in multi-tenant scenarios, one tenant's data can occupy node cache for extended periods. This leaves insufficient cache capacity for the next tenant, who is then forced to fall back to object storage, causing a noticeable performance drop.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;To address this, we deployed an independent daemon process on each node that performs a global cache garbage collection (GC) pass before the application FUSE client starts. The eviction strategy references the JuiceFS FUSE client implementation, using a 2-random policy to balance collection efficiency and performance overhead. Coordination across nodes is handled via Kubernetes distributed locks: only the client that acquires the lock executes GC, preventing multiple clients from running cache collection simultaneously and creating excessive network and I/O pressure.&lt;/p&gt;

&lt;p&gt;This mechanism effectively mitigates the problem of historical jobs occupying cache resources in multi-tenant scenarios, allowing different tenants sharing elastic resources to maintain consistent cache performance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;For elastic compute to reliably serve production traffic, compute scheduling alone is not enough. Model data and hot data must remain stably accessible across clouds and clusters.&lt;/p&gt;

&lt;p&gt;Built on JuiceFS, we’ve combined object storage, unified namespace, metadata management, distributed caching, and FUSE mounting into an object storage acceleration solution purpose-built for elastic inference. This is not simply about mounting object storage as a file system. It’s about building a data access layer around the access patterns of model inference: one that supports warm-up, caching, isolation, and management.&lt;/p&gt;

&lt;p&gt;This represents Gongjiyun's current progress in elastic compute and cross-cloud storage acceleration. As AI inference scenarios continue to evolve, model distribution, cache management, and multi-cluster data access will continue to surface new engineering challenges. We look forward to exchanging ideas with developers, AI application teams, and infrastructure practitioners, and to exploring more stable and efficient data access solutions for elastic compute environments.&lt;/p&gt;

&lt;p&gt;If you have any questions for this article, feel free to join &lt;a href="https://github.com/juicedata/juicefs/discussions/" rel="noopener noreferrer"&gt;JuiceFS discussions on GitHub&lt;/a&gt; and &lt;a href="http://go.juicefs.com/discord" rel="noopener noreferrer"&gt;community on Discord&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
  </channel>
</rss>
