DEV Community

john wong
john wong

Posted on

I Bricked My DGX Spark's Inference Engine With a git pull. Here's How I Got It Back.

I run an Nvidia DGX Spark — the little GB10 ARM box with 128 GB of unified memory. My morning ritual for the past week had been simple:

cd spark-vllm-docker
git pull
./run-recipe.sh qwen3.8-flash-next-nvfp4-solo --solo --earlyoom --setup
Enter fullscreen mode Exit fullscreen mode

Pull the recipe repo, rebuild, serve. This morning it stopped being simple.

The failure

Mid-startup, the engine died:

(EngineCore pid=199) ERROR 10-03 00:41:33 [core.py:1483] EngineCore failed to start.
(EngineCore pid=199) ERROR 10-03 00:41:33 [core.py:1483] Traceback (most recent call last):
(EngineCore pid=199) ERROR ... File ".../vllm/v1/engine/core.py", line 1445, in run_engine_core
Enter fullscreen mode Exit fullscreen mode

Every attempt after that — fresh or with tweaks — ended with the same final line:

RuntimeError: Engine core initialization failed. See root cause above. Failed core proc(s):
Stopping cluster...
Cluster stopped.
Enter fullscreen mode Exit fullscreen mode

A git pull had moved the recipe onto a new vLLM release, and that release is broken for this model on a single Spark. The repo's own history couldn't save me either:

$ git checkout 71c26b8
error: pathspec '71c26b8' did not match any file(s) known to git
Enter fullscreen mode Exit fullscreen mode

(The commit wasn't in my clone's history. Lesson one: don't assume you can roll back a shallow/partial clone.)

The two hours of things that could not possibly work

Here is where I made it worse, and it's the most useful part of this story.

I asked chatbots. Confidently, they told me to "downgrade to the stable V0 engine":

./run-recipe.sh recipes/<your-recipe>.yaml -- --vllm-v1-disable
Enter fullscreen mode Exit fullscreen mode

That flag does not exist. The V0 engine was fully removed from vLLM (the VLLM_USE_V1 variable was deleted from the codebase). There is no V0 to switch back to. Another agent told me to hand-edit sandboxes.json — which corrupted my gateway config and bought me a second incident on top of the first.

I burned real time on plausible-sounding flags: --enforce-eager, PYTORCH_CUDA_ALLOC_CONF=expandable_segments:False, --disable-custom-all-reduce. All failed, for a good reason: when the breakage is engine-wide (a code release, not a knob), no flag fixes it.

And a detail the terminal taught me that no agent mentioned: half my "Hugging Face gated repo" problem was actually a file-permission problem. The weights were on disk the whole time — as root-owned files the container couldn't read:

PermissionError: [Errno 13] Permission denied:
 '.../models--local-inference-lab--Qwen3.8-Flash-Next-NVFP4/snapshots/.../model-00024-of-00041.safetensors'
$ sudo chown -R skyspark:skyspark ~/.cache/huggingface
Enter fullscreen mode Exit fullscreen mode

Before chasing an HTTP 403, run ls -l on the cache dir. chown is frequently the whole fix.

Lesson two: a chat agent will invent a flag with total confidence. Before running any command it suggests, check the tool's real docs or --help. The fastest way to diagnose the EngineCore crash was the one neither agent volunteered: git log on the vLLM release, which showed a known-broken version had landed.

The real math problem

The model I serve, Qwen3.8-Flash-Next (NVFP4 quant), has a ~124 GiB checkpoint. The Spark's unified pool is 128 GB. With KV cache and everything else, it simply doesn't fit the way vanilla vLLM wants to load it — 48 GiB of that checkpoint is an n-gram embedding lookup table that any given token only touches ~16 rows of.

The fix that works is architectural, not a flag: patch vLLM to serve the lookup table from NVMe via mmap instead of pinning it in memory. Resident weights drop to ~75 GiB, the rest of the pool goes to KV.

The community recipe that does exactly this, patched against vLLM v0.30.0 with the mmap trick plus two GB10 bug fixes:

git clone https://github.com/blazux/qwen3.8-Flash-DGX.git && cd qwen3.8-Flash-DGX
./flash doctor    # sanity: docker, GPU, memory, disk, port, image, weights
./flash setup     # build image, download checkpoint (resumable), prep hybrid layout
./flash serve     # recommended profile: hybrid, 500k ctx, deterministic
./flash wait      # first boot loads ~75 GiB of weights, 3-4 min
time ./flash test # health, coherence, prefix-cache, determinism, tok/s
Enter fullscreen mode Exit fullscreen mode

My engine was back on http://192.168.0.52:18300/v1, with an OpenAI-compatible endpoint and the same model name. What confirmed the patched build was actually running, in the startup log:

(EngineCore pid=362) INFO PLE mmap patch applied to ...Qwen4ExpNGramEmbedding
(EngineCore pid=362) INFO fp8 hybrid (modelopt): 1836 blockwise-fp8 layers detected
INFO: Application startup complete.
INFO: "GET /v1/models HTTP/1.1" 200 OK
(EngineCore pid=362) INFO PLE mmap stats: 9 ops, 23 ms total, 0.0 MiB read, gpu-wait 0.08 ms/op
Enter fullscreen mode Exit fullscreen mode

No PLE-mmap line in your log? You're running the unpatched image — which is precisely the broken state to begin with.

The second bug: "my sandbox can't reach the engine"

Fixed the engine, and the agent container couldn't see it:

inference.local returned transient HTTP 503; response_bytes=41
...
Validation probe summary: SSRF preflight: no HTTP response.
Enter fullscreen mode Exit fullscreen mode

The onboarding tool refuses the Docker bridge gateway address (172.18.0.1) by default — an SSRF guard, not a bug. Marking it trusted at onboarding time solved it in one line:

NEMOCLAW_TRUSTED_PRIVATE_HOSTS=172.18.0.1 nemoclaw onboard --fresh --name spark4
Enter fullscreen mode Exit fullscreen mode

On the PC side, the client needed the LAN address for the same engine, and the new port: base_url: http://192.168.0.52:18300/v1. One service, three addresses (localhost:18300, 172.18.0.1:18300, 192.168.0.52:18300) — every layer of your stack is a different network.

What I'd do differently

  1. Check the changelog before the pull. A recipe git pull that silently moves you onto a new vLLM release is a release-upgrade decision, not a no-op sync. Pin the recipe, review what changed, then merge.
  2. Never hand-edit state files an onboarding tool owns — sandboxes.json and friends. Redo the supported command with the right flag instead.
  3. Treat agent-suggested flags as hypotheses, not commands. --vllm-v1-disable was pure fiction.
  4. Screenshot your terminal during an incident. Half this post was reconstructed from them — the crash trace and git pathspec error would otherwise be memory.
  5. ./flash doctor + ./flash test after every infra event (reboot, network switch, image update). The nasty failures on this box are the ones that leave a health endpoint green and the quality quietly wrong.

The Spark is a weird, wonderful machine — a desktop with server-class unified memory. It punishes you the moment you treat its memory as infinite. But the fix was never a flag: it was knowing which 48 GiB you can stop paying for.

Three more things the screenshots taught me

Since publishing the first draft I OCR'd my incident screenshots — more terminal history than
memory — and found three extra lessons:

  1. The container port map matters. docker ps showed the engine publishing 0.0.0.0:18300->8000 — vLLM listens on 8000 inside, the host maps 18300. A stale sandbox config probing host:8000 fails with curl (7) and looks exactly like "the engine is down" when it is up and healthy. Grep every layer's config for the old port.
  2. FP8 KV cache is a quality trade, not a free win. One recipe's memory-budget report warned that fp8 KV gives ~1.7x more KV tokens (1M context looks reachable) but dropped a long-reasoning benchmark from 6/6 to 2/6 — with sparse attention, quantized keys perturb which blocks the indexer selects. Re-validate quality on your own workload before trading it for context length.
  3. Your host kernel may be mis-tuned for the NVIDIA driver. The same report flagged vm.min_free_kbytes=45166 and vm.watermark_scale_factor=10 at defaults — no free-page reserve for the GPU driver. The recipe won't apply it for you (sudo sysctl -p is a one-time host step). It's recipe-agnostic: check it once, keep it forever.

Engine: patched vLLM v0.30.0 (blazux/qwen3.8-Flash-DGX, Apache-2.0) on a single DGX Spark. Checkpoint: nvidia/Qwen3.8-Flash-Next-NVFP4.

Top comments (0)