DEV Community

Francisco Booth
Francisco Booth

Posted on

Why HF_HOME Points To Ephemeral Storage On RunPod By Default And How To Fix It

If you have ever run a HuggingFace model download on RunPod and come back to find your cache gone, this is why.


The default is wrong for rented GPU pods

HuggingFace sets HF_HOME to ~/.cache/huggingface when no environment variable overrides it. On a standard Linux machine that is fine your home directory persists.

On RunPod, your home directory is on the container volume, which is ephemeral. When the pod stops, it is gone.

RunPod has two storage layers: the container volume, which is ephemeral, and the network volume mounted at /workspace, which persists across pod restarts. HF_HOME defaults to the wrong one.

This means every model download, every tokeniser, every dataset you pulled via datasets.load_dataset() gone on pod restart. You pay to download them again on the next run.

On large models that is not a minor inconvenience. A 7B model download at pod startup adds minutes and real cost every single time.

The fix is one line in your startup script:


export HF_HOME=/workspace/.cache/huggingface


/workspace is RunPod’s persistent volume mount. One line and your cache survives every restart.

Why this is easy to miss

There is no error message. HuggingFace downloads the files successfully to the ephemeral path. Your training run completes. The problem only surfaces when you restart the pod and the download happens again and even then it looks like a slow startup, not a misconfiguration.

Silent failures are the hardest to debug because you are not looking for them.

Automating the check

If you want this caught automatically before every job rather than relying on remembering to set the variable, ComputeFence checks it as part of its pre-flight scan:


uvx computefence doctor


It flags HF_HOME pointing to ephemeral storage and prints the exact export command for your platform. It also checks GPU visibility, Accelerate config, checkpoint directory persistence, and disk headroom the other silent failures that cost compute budget without a log entry.

The two-line pre-flight for every RunPod job


export HF_HOME=/workspace/.cache/huggingface
uvx computefence doctor


The first line fixes the cache. The second confirms everything else is configured correctly before the GPU bill starts.

GitHub: github.com/Francisco-Booth/ComputeFence
Install: pip install computefence

Top comments (0)