I run an Nvidia DGX Spark — the little GB10 ARM box with 128 GB of unified memory. My morning ritual for the past week had been simple:
cd spark-vllm-docker
git pull
./run-recipe.sh qwen3.8-flash-next-nvfp4-solo --solo --earlyoom --setup
Pull the recipe repo, rebuild, serve. This morning it stopped being simple.
The failure
Mid-startup, the engine died:
(EngineCore pid=199) ERROR 10-03 00:41:33 [core.py:1483] EngineCore failed to start.
(EngineCore pid=199) ERROR 10-03 00:41:33 [core.py:1483] Traceback (most recent call last):
(EngineCore pid=199) ERROR ... File ".../vllm/v1/engine/core.py", line 1445, in run_engine_core
Every attempt after that — fresh or with tweaks — ended with the same final line:
RuntimeError: Engine core initialization failed. See root cause above. Failed core proc(s):
Stopping cluster...
Cluster stopped.
A git pull had moved the recipe onto a new vLLM release, and that release is broken for this model on a single Spark. The repo's own history couldn't save me either:
$ git checkout 71c26b8
error: pathspec '71c26b8' did not match any file(s) known to git
(The commit wasn't in my clone's history. Lesson one: don't assume you can roll back a shallow/partial clone.)
The two hours of things that could not possibly work
Here is where I made it worse, and it's the most useful part of this story.
I asked chatbots. Confidently, they told me to "downgrade to the stable V0 engine":
./run-recipe.sh recipes/<your-recipe>.yaml -- --vllm-v1-disable
That flag does not exist. The V0 engine was fully removed from vLLM (the VLLM_USE_V1 variable was deleted from the codebase). There is no V0 to switch back to. Another agent told me to hand-edit sandboxes.json — which corrupted my gateway config and bought me a second incident on top of the first.
I burned real time on plausible-sounding flags: --enforce-eager, PYTORCH_CUDA_ALLOC_CONF=expandable_segments:False, --disable-custom-all-reduce. All failed, for a good reason: when the breakage is engine-wide (a code release, not a knob), no flag fixes it.
And a detail the terminal taught me that no agent mentioned: half my "Hugging Face gated repo" problem was actually a file-permission problem. The weights were on disk the whole time — as root-owned files the container couldn't read:
PermissionError: [Errno 13] Permission denied:
'.../models--local-inference-lab--Qwen3.8-Flash-Next-NVFP4/snapshots/.../model-00024-of-00041.safetensors'
$ sudo chown -R skyspark:skyspark ~/.cache/huggingface
Before chasing an HTTP 403, run ls -l on the cache dir. chown is frequently the whole fix.
Lesson two: a chat agent will invent a flag with total confidence. Before running any command it suggests, check the tool's real docs or
--help. The fastest way to diagnose the EngineCore crash was the one neither agent volunteered:git logon the vLLM release, which showed a known-broken version had landed.
The real math problem
The model I serve, Qwen3.8-Flash-Next (NVFP4 quant), has a ~124 GiB checkpoint. The Spark's unified pool is 128 GB. With KV cache and everything else, it simply doesn't fit the way vanilla vLLM wants to load it — 48 GiB of that checkpoint is an n-gram embedding lookup table that any given token only touches ~16 rows of.
The fix that works is architectural, not a flag: patch vLLM to serve the lookup table from NVMe via mmap instead of pinning it in memory. Resident weights drop to ~75 GiB, the rest of the pool goes to KV.
The community recipe that does exactly this, patched against vLLM v0.30.0 with the mmap trick plus two GB10 bug fixes:
git clone https://github.com/blazux/qwen3.8-Flash-DGX.git && cd qwen3.8-Flash-DGX
./flash doctor # sanity: docker, GPU, memory, disk, port, image, weights
./flash setup # build image, download checkpoint (resumable), prep hybrid layout
./flash serve # recommended profile: hybrid, 500k ctx, deterministic
./flash wait # first boot loads ~75 GiB of weights, 3-4 min
time ./flash test # health, coherence, prefix-cache, determinism, tok/s
My engine was back on http://192.168.0.52:18300/v1, with an OpenAI-compatible endpoint and the same model name. What confirmed the patched build was actually running, in the startup log:
(EngineCore pid=362) INFO PLE mmap patch applied to ...Qwen4ExpNGramEmbedding
(EngineCore pid=362) INFO fp8 hybrid (modelopt): 1836 blockwise-fp8 layers detected
INFO: Application startup complete.
INFO: "GET /v1/models HTTP/1.1" 200 OK
(EngineCore pid=362) INFO PLE mmap stats: 9 ops, 23 ms total, 0.0 MiB read, gpu-wait 0.08 ms/op
No PLE-mmap line in your log? You're running the unpatched image — which is precisely the broken state to begin with.
The second bug: "my sandbox can't reach the engine"
Fixed the engine, and the agent container couldn't see it:
inference.local returned transient HTTP 503; response_bytes=41
...
Validation probe summary: SSRF preflight: no HTTP response.
The onboarding tool refuses the Docker bridge gateway address (172.18.0.1) by default — an SSRF guard, not a bug. Marking it trusted at onboarding time solved it in one line:
NEMOCLAW_TRUSTED_PRIVATE_HOSTS=172.18.0.1 nemoclaw onboard --fresh --name spark4
On the PC side, the client needed the LAN address for the same engine, and the new port: base_url: http://192.168.0.52:18300/v1. One service, three addresses (localhost:18300, 172.18.0.1:18300, 192.168.0.52:18300) — every layer of your stack is a different network.
What I'd do differently
-
Check the changelog before the pull. A recipe
git pullthat silently moves you onto a new vLLM release is a release-upgrade decision, not a no-op sync. Pin the recipe, review what changed, then merge. -
Never hand-edit state files an onboarding tool owns —
sandboxes.jsonand friends. Redo the supported command with the right flag instead. -
Treat agent-suggested flags as hypotheses, not commands.
--vllm-v1-disablewas pure fiction. -
Screenshot your terminal during an incident. Half this post was reconstructed from them — the crash trace and
git pathspecerror would otherwise be memory. -
./flash doctor+./flash testafter every infra event (reboot, network switch, image update). The nasty failures on this box are the ones that leave a health endpoint green and the quality quietly wrong.
The Spark is a weird, wonderful machine — a desktop with server-class unified memory. It punishes you the moment you treat its memory as infinite. But the fix was never a flag: it was knowing which 48 GiB you can stop paying for.
Three more things the screenshots taught me
Since publishing the first draft I OCR'd my incident screenshots — more terminal history than
memory — and found three extra lessons:
-
The container port map matters.
docker psshowed the engine publishing0.0.0.0:18300->8000— vLLM listens on 8000 inside, the host maps 18300. A stale sandbox config probinghost:8000fails withcurl (7)and looks exactly like "the engine is down" when it is up and healthy. Grep every layer's config for the old port. - FP8 KV cache is a quality trade, not a free win. One recipe's memory-budget report warned that fp8 KV gives ~1.7x more KV tokens (1M context looks reachable) but dropped a long-reasoning benchmark from 6/6 to 2/6 — with sparse attention, quantized keys perturb which blocks the indexer selects. Re-validate quality on your own workload before trading it for context length.
-
Your host kernel may be mis-tuned for the NVIDIA driver. The same report flagged
vm.min_free_kbytes=45166andvm.watermark_scale_factor=10at defaults — no free-page reserve for the GPU driver. The recipe won't apply it for you (sudo sysctl -pis a one-time host step). It's recipe-agnostic: check it once, keep it forever.
Engine: patched vLLM v0.30.0 (blazux/qwen3.8-Flash-DGX, Apache-2.0) on a single DGX Spark. Checkpoint: nvidia/Qwen3.8-Flash-Next-NVFP4.
Top comments (0)