DEV Community

ZAKARIA KHCHICHE
ZAKARIA KHCHICHE

Posted on Originally published at Medium

Running Mistral offline: the tokenizer is on disk, but the library can't find it

An inference server cut off from the Internet. A dedicated disk for models. Mistral's tokenizer downloaded in advance, sitting in its folder. And at startup:

FileNotFoundError: No local files found for the repo ID mistralai/Mistral-7B-v0.1 and revision None.
Enter fullscreen mode Exit fullscreen mode

The file is there. The library looks somewhere else.

I found this bug in mistral-common, Mistral AI's open-source library that prepares requests for its models. The fix was reviewed, approved and merged by a Mistral AI maintainer (PR #349, issue #348).

Who is affected

Teams running Mistral models without Internet access that store their models outside the default Hugging Face cache. That is exactly the enterprise setup: air-gapped network for security, corporate proxy, a dedicated model disk on the GPU server.

Two situations trigger it:

  • the caller forbids network access (local_files_only=True);
  • the Hub is unreachable (outage, firewall, HF_HUB_OFFLINE=1) and the library falls back to local files on its own.

With the default cache, or with network access, nothing happens. That is why the bug stays invisible in development and shows up on deployment day.

The cause: two steps, two folders

Offline, loading the tokenizer chains two functions:

  1. list_local_hf_repo_files lists the cached files to find the tokenizer (tekken.json). It always looked in the default cache, HF_HUB_CACHE, and had no cache_dir parameter.
  2. hf_hub_download then fetches the file. It does receive cache_dir and looks in the custom folder.
# before: the listing ignores the folder chosen by the user
repo_cache = Path(huggingface_hub.constants.HF_HUB_CACHE) / ...
Enter fullscreen mode Exit fullscreen mode

If the tokenizer only lives in cache_dir, the listing comes back empty and the function gives up before even trying to read the file. The reverse case fails too, one step later.

The fix: 5 lines, the same rule as Hugging Face

Both steps now read the same folder: cache_dir if given, otherwise the default cache, exactly the rule hf_hub_download follows.

cache_root = Path(cache_dir) if cache_dir is not None else Path(huggingface_hub.constants.HF_HUB_CACHE)
repo_cache = cache_root / huggingface_hub.constants.REPO_ID_SEPARATOR.join(["models", *repo_id.split("/")])
Enter fullscreen mode Exit fullscreen mode
  • Backward compatible: a new optional parameter, unchanged behavior without it.
  • 6 regression tests, 5 of which fail without the fix: folder given as a string or a Path, full offline load, fallback after a network error, HF_HUB_OFFLINE=1.
  • Minimal scope: no new feature slipped into the fix.

The maintainer asked for a single change and applied it himself: removing a code comment. The code itself did not change.

Until you can upgrade

If you are affected and cannot upgrade mistral-common yet: point HF_HUB_CACHE at the same folder as your cache_dir. Both steps then read the same place.

From reading their source code, Transformers (cache_dir) and vLLM (--download-dir with tokenizer_mode=mistral) pass that folder down to mistral-common, so they are probably affected offline too. I did not test that end to end.

What this means for teams shipping agents

  1. Test the real deployment, not the developer laptop. This bug only exists offline with a custom model folder: exactly the production setup, never the laptop one.
  2. Cut the network in your tests. HF_HUB_OFFLINE=1 in one integration test would have caught it.
  3. An option that is not passed everywhere is a bug. When a function accepts cache_dir, check that every internal call receives it.
  4. Contribute back. Five lines and six tests is the price of a fix that protects everyone deploying Mistral offline.

Deploying Mistral models on an air-gapped network? Which other traps have you hit?

Let's talk, or get your own agents audited: LinkedIn · Malt · Medium · Website

Want your team to build agents like these? I run a hands-on Copilot Studio training (in French, with Spar-x, Qualiopi-certified, eligible for OPCO funding in France): https://zakariakhchiche.github.io/formation-copilot-studio/ and a free AI Act article 4 kit: https://zakariakhchiche.github.io/kit-ai-act/

Top comments (0)