A question we keep getting from people trying HuskHoard for the first time is some version of: "So it archives to S3... but how is that different from just writing a script that uploads old files to a bucket?" It's a fair question, and the honest answer is that the difference isn't really about the upload. It's about what happens after the upload — how the file still behaves like a normal file on your disk, and how HuskHoard gets exactly the bytes it needs back out of the cloud without downloading the world.
This post walks through that whole loop: how a file gets turned into a stub, how the payload actually lands in a cloud bucket, and what happens on the way back when someone opens it again.
The Core Idea: The File System Is a Presentation Layer
Most of us grow up assuming that a file path is the physical truth — that project_final.mov sitting in a folder means those exact bytes are sitting on the disk under that folder. HuskHoard breaks that assumption on purpose.
Mainframe shops have known for decades that you can separate the namespace (the folder structure your OS and applications see) from the storage tier (where the bytes physically live). That's Hierarchical Storage Management, and HuskHoard brings the same idea to a normal Linux box, written in Rust, with no proprietary format sitting between you and your data.
In practice, this means a file can look completely normal — correct size, correct timestamps, correct permissions — while occupying next to nothing on your fast local drive. That's a stub, and it's the mechanism that makes everything else possible.
Turning a File Into a Stub
When the Janitor (HuskHoard's background policy daemon) decides a file is cold — based on age, extension, or directory rules you configure — it copies the payload off to the archive tier and then hollows out the local copy. It doesn't do this with a symlink or a shortcut. It uses the Linux fallocate syscall with FALLOC_FL_PUNCH_HOLE, which deallocates the underlying data blocks while leaving the file's apparent size untouched.
The practical result is that ls -l still reports the original size, but du -sh drops to essentially nothing. In practice that "nothing" is a single filesystem block — typically 4 KB, the minimum an inode can allocate — since the residual xattrs and metadata still have to live somewhere. So a stubbed 50 GB file won't show as a literal 0 B; it'll floor out at 4 KB (or whatever your block size is) regardless of how large the original payload was. HuskHoard then tags the file with an extended attribute (trusted.husk.status = stubbed) and quietly restores the original modification and access timestamps with utimensat, so nothing downstream — your backup tool, your IDE, your indexing service — has any reason to notice the file was touched.
None of this would be usable without a ground truth pointing back to where the real bytes went, which is where the Catalog comes in: a local SQLite database (in WAL mode) that records, per file, the compression ratio, a BLAKE3 checksum, and the exact volume and byte offset the payload lives at. Lose your hot-tier SSD entirely, and as long as the Catalog and the archive survive, HuskHoard can regenerate every stub and rebuild your directory tree in seconds.
Getting Files Into the Cloud
This is where HuskHoard treats cloud storage as, essentially, a very fast tape drive. To the archiving engine, S3, Backblaze B2, and similar services are just high-latency, append-only mediums — conceptually the same as an LTO cartridge, just reached over HTTP instead of a SCSI bus. Because HuskHoard shells out to Rclone for the actual transport, any of the dozens of providers Rclone supports can sit behind a Husk volume without HuskHoard needing a bespoke integration for each one.
The part that actually matters for your cloud bill is how files get packed before they're sent. If HuskHoard uploaded a one-to-one mirror of your filesystem, archiving something like a node_modules directory would generate an enormous number of PUT requests, and most providers charge per API call, not just per gigabyte. So instead, the Archive Worker streams files through Zstd compression and BLAKE3 hashing and pipes the result into large, sequentially written objects — husk_4096.bin, husk_1610612736.bin, and so on — often 100+ GB each. Your bucket ends up holding a handful of big binary blobs rather than a mirror of your directory tree. The Catalog is what remembers which byte range inside which .bin object corresponds to which original file.
Within those objects, payloads are further chunked into 16 MB "Jump Frames," each with a small TLV index at the front so a specific frame can be located without decompressing everything before it. That indexing step is what makes selective retrieval possible later — without it, pulling a single document out of a 150 GB archive object would mean streaming and discarding everything before it just to rebuild decompression state.
Getting Files Back Out
Here's the part that's genuinely different from a manual bucket-based workflow. If you're managing cold storage by hand — say, syncing old files to S3 and deleting the local copies — retrieving one usually means either remembering it's gone and running a script, or building your own placeholder/rehydration logic. HuskHoard makes that step invisible.
When an application calls open() on a stubbed file, HuskHoard intercepts the request using fanotify, a Linux kernel API that lets a userspace process pause another process's file access before it completes. The kernel blocks the requesting application; HuskHoard checks the file's xattr, sees it's stubbed, and queries the Catalog for the exact volume and offset the payload lives at.
If that volume is a cloud bucket, HuskHoard's StreamGate component spawns a targeted Rclone call — something like rclone cat husk_4096.bin --offset X --count Y — which Rclone translates into an HTTP Range request. The cloud provider serves back only the specific slice of bytes needed, not the entire multi-gigabyte object. HuskHoard decompresses that slice in memory, verifies it against the BLAKE3 checksum, streams it into the now-un-punched file, strips the "stubbed" xattr, and releases the paused process. From the application's point of view, the read just took a little longer than usual — like a slow disk, not a missing file.
That range-request behavior is the actual source of the cost savings people notice. Pulling a 5 MB PDF out of a 150 GB archive object costs 5 MB of egress, not 150 GB. For already-compressed media like MP4 or MOV, the byte offsets line up close to 1:1, so seeking through a video file streamed from a bucket is close to instantaneous, with no need to decompress anything preceding the requested section.
Two Ways to Deploy It
Because the storage backend is abstracted this way, there are two reasonable places to put HuskHoard itself:
Local Gateway. HuskHoard runs on a machine in your office or homelab. Your hot tier is a local NVMe drive or ZFS array; your cold tier is a cloud bucket. You get LAN-speed access to active work, and cold data quietly drains off to the cloud on a schedule, with the local footprint shrinking down to stub-only sizes. This suits teams working with heavy local assets — video, renders, large datasets — where cloud archiving is essentially a hands-off offsite backup.
Cloud-Native Hub. HuskHoard runs on a cloud compute instance instead, with a small, fast block volume (e.g., AWS EBS) as the hot tier and an object store as the cold tier. Remote users connect over VPN, SMB, or HuskHoard's built-in HTTP gateway. As the hot volume fills, HuskHoard spills the oldest files out to the object store automatically, so a 500 GB EBS volume can present as a 500 TB share to everyone connecting to it. This fits distributed teams working on transactional, document-heavy workloads where there's no single office to anchor local storage to.
Why This Beats Managing Buckets Manually
A hand-rolled "sync old files to S3" setup can move data, but it usually can't do three things HuskHoard does by default: it can't make the cloud copy look and behave like a normal, present file to every application on the box; it can't retrieve a fragment of a file without pulling the whole object; and it doesn't give you a portable, provider-agnostic map of what went where. Because HuskHoard's payload format is just Zstd streams and BLAKE3 checksums — no proprietary container — you can extract everything with standard tools even if you stopped using HuskHoard tomorrow. Swapping providers is a config change, not a migration project, since the abstraction sits on top of Rclone rather than any single vendor's SDK.
That's the actual pitch: not that HuskHoard uploads files to the cloud, but that it makes "the cloud" disappear as a concept your applications, or your teammates, ever have to think about.
Top comments (0)