<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: john wong</title>
    <description>The latest articles on DEV Community by john wong (@john_wong_716289005e1fa4e).</description>
    <link>https://dev.to/john_wong_716289005e1fa4e</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4159880%2F9430b21d-32ad-411b-a0fe-ac6620573bb1.jpg</url>
      <title>DEV Community: john wong</title>
      <link>https://dev.to/john_wong_716289005e1fa4e</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/john_wong_716289005e1fa4e"/>
    <language>en</language>
    <item>
      <title>I Bricked My DGX Spark's Inference Engine With a git pull. Here's How I Got It Back.</title>
      <dc:creator>john wong</dc:creator>
      <pubDate>Sat, 03 Oct 2026 17:44:40 +0000</pubDate>
      <link>https://dev.to/john_wong_716289005e1fa4e/i-bricked-my-dgx-sparks-inference-engine-with-a-git-pull-heres-how-i-got-it-back-47n1</link>
      <guid>https://dev.to/john_wong_716289005e1fa4e/i-bricked-my-dgx-sparks-inference-engine-with-a-git-pull-heres-how-i-got-it-back-47n1</guid>
      <description>&lt;p&gt;I run an Nvidia DGX Spark — the little GB10 ARM box with 128 GB of unified memory. My morning ritual for the past week had been simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd &lt;/span&gt;spark-vllm-docker
git pull
./run-recipe.sh qwen3.8-flash-next-nvfp4-solo &lt;span class="nt"&gt;--solo&lt;/span&gt; &lt;span class="nt"&gt;--earlyoom&lt;/span&gt; &lt;span class="nt"&gt;--setup&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pull the recipe repo, rebuild, serve. This morning it stopped being simple.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure
&lt;/h2&gt;

&lt;p&gt;Mid-startup, the engine died:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(EngineCore pid=199) ERROR 10-03 00:41:33 [core.py:1483] EngineCore failed to start.
(EngineCore pid=199) ERROR 10-03 00:41:33 [core.py:1483] Traceback (most recent call last):
(EngineCore pid=199) ERROR ... File ".../vllm/v1/engine/core.py", line 1445, in run_engine_core
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every attempt after that — fresh or with tweaks — ended with the same final line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;RuntimeError: Engine core initialization failed. See root cause above. Failed core proc(s):
Stopping cluster...
Cluster stopped.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;code&gt;git pull&lt;/code&gt; had moved the recipe onto a new vLLM release, and that release is broken for this model on a single Spark. The repo's own history couldn't save me either:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$ git checkout 71c26b8
error: pathspec '71c26b8' did not match any file(s) known to git
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(The commit wasn't in my clone's history. Lesson one: don't assume you can roll back a shallow/partial clone.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The two hours of things that could not possibly work
&lt;/h2&gt;

&lt;p&gt;Here is where I made it worse, and it's the most useful part of this story.&lt;/p&gt;

&lt;p&gt;I asked chatbots. Confidently, they told me to "downgrade to the stable V0 engine":&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;./run-recipe.sh recipes/&amp;lt;your-recipe&amp;gt;.yaml &lt;span class="nt"&gt;--&lt;/span&gt; &lt;span class="nt"&gt;--vllm-v1-disable&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;That flag does not exist.&lt;/strong&gt; The V0 engine was fully removed from vLLM (the &lt;code&gt;VLLM_USE_V1&lt;/code&gt; variable was deleted from the codebase). There is no V0 to switch back to. Another agent told me to hand-edit &lt;code&gt;sandboxes.json&lt;/code&gt; — which corrupted my gateway config and bought me a second incident on top of the first.&lt;/p&gt;

&lt;p&gt;I burned real time on plausible-sounding flags: &lt;code&gt;--enforce-eager&lt;/code&gt;, &lt;code&gt;PYTORCH_CUDA_ALLOC_CONF=expandable_segments:False&lt;/code&gt;, &lt;code&gt;--disable-custom-all-reduce&lt;/code&gt;. All failed, for a good reason: when the breakage is engine-wide (a code release, not a knob), no flag fixes it.&lt;/p&gt;

&lt;p&gt;And a detail the terminal taught me that no agent mentioned: half my "Hugging Face gated repo" problem was actually a &lt;strong&gt;file-permission problem&lt;/strong&gt;. The weights were on disk the whole time — as root-owned files the container couldn't read:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PermissionError: [Errno 13] Permission denied:
 '.../models--local-inference-lab--Qwen3.8-Flash-Next-NVFP4/snapshots/.../model-00024-of-00041.safetensors'
$ sudo chown -R skyspark:skyspark ~/.cache/huggingface
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Before chasing an HTTP 403, run &lt;code&gt;ls -l&lt;/code&gt; on the cache dir. &lt;code&gt;chown&lt;/code&gt; is frequently the whole fix.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Lesson two: a chat agent will invent a flag with total confidence. Before running any command it suggests, check the tool's real docs or &lt;code&gt;--help&lt;/code&gt;.&lt;/strong&gt; The fastest way to diagnose the EngineCore crash was the one neither agent volunteered: &lt;code&gt;git log&lt;/code&gt; on the vLLM release, which showed a known-broken version had landed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The real math problem
&lt;/h2&gt;

&lt;p&gt;The model I serve, Qwen3.8-Flash-Next (NVFP4 quant), has a ~124 GiB checkpoint. The Spark's unified pool is 128 GB. With KV cache and everything else, it simply doesn't fit the way vanilla vLLM wants to load it — 48 GiB of that checkpoint is an n-gram embedding lookup table that any given token only touches ~16 rows of.&lt;/p&gt;

&lt;p&gt;The fix that works is architectural, not a flag: patch vLLM to &lt;strong&gt;serve the lookup table from NVMe via &lt;code&gt;mmap&lt;/code&gt;&lt;/strong&gt; instead of pinning it in memory. Resident weights drop to ~75 GiB, the rest of the pool goes to KV.&lt;/p&gt;

&lt;p&gt;The community recipe that does exactly this, patched against vLLM v0.30.0 with the mmap trick plus two GB10 bug fixes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/blazux/qwen3.8-Flash-DGX.git &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;cd &lt;/span&gt;qwen3.8-Flash-DGX
./flash doctor    &lt;span class="c"&gt;# sanity: docker, GPU, memory, disk, port, image, weights&lt;/span&gt;
./flash setup     &lt;span class="c"&gt;# build image, download checkpoint (resumable), prep hybrid layout&lt;/span&gt;
./flash serve     &lt;span class="c"&gt;# recommended profile: hybrid, 500k ctx, deterministic&lt;/span&gt;
./flash &lt;span class="nb"&gt;wait&lt;/span&gt;      &lt;span class="c"&gt;# first boot loads ~75 GiB of weights, 3-4 min&lt;/span&gt;
&lt;span class="nb"&gt;time&lt;/span&gt; ./flash &lt;span class="nb"&gt;test&lt;/span&gt; &lt;span class="c"&gt;# health, coherence, prefix-cache, determinism, tok/s&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;My engine was back on &lt;code&gt;http://192.168.0.52:18300/v1&lt;/code&gt;, with an OpenAI-compatible endpoint and the same model name. What confirmed the patched build was actually running, in the startup log:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;(EngineCore pid=362) INFO PLE mmap patch applied to ...Qwen4ExpNGramEmbedding
(EngineCore pid=362) INFO fp8 hybrid (modelopt): 1836 blockwise-fp8 layers detected
INFO: Application startup complete.
INFO: "GET /v1/models HTTP/1.1" 200 OK
(EngineCore pid=362) INFO PLE mmap stats: 9 ops, 23 ms total, 0.0 MiB read, gpu-wait 0.08 ms/op
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No PLE-mmap line in your log? You're running the unpatched image — which is precisely the broken state to begin with.&lt;/p&gt;

&lt;h2&gt;
  
  
  The second bug: "my sandbox can't reach the engine"
&lt;/h2&gt;

&lt;p&gt;Fixed the engine, and the agent container couldn't see it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;inference.local returned transient HTTP 503; response_bytes=41
...
Validation probe summary: SSRF preflight: no HTTP response.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The onboarding tool refuses the Docker bridge gateway address (&lt;code&gt;172.18.0.1&lt;/code&gt;) by default — an SSRF guard, not a bug. Marking it trusted at onboarding time solved it in one line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;NEMOCLAW_TRUSTED_PRIVATE_HOSTS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;172.18.0.1 nemoclaw onboard &lt;span class="nt"&gt;--fresh&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; spark4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On the PC side, the client needed the LAN address for the same engine, and the new port: &lt;code&gt;base_url: http://192.168.0.52:18300/v1&lt;/code&gt;. One service, three addresses (&lt;code&gt;localhost:18300&lt;/code&gt;, &lt;code&gt;172.18.0.1:18300&lt;/code&gt;, &lt;code&gt;192.168.0.52:18300&lt;/code&gt;) — every layer of your stack is a different network.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd do differently
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Check the changelog before the pull.&lt;/strong&gt; A recipe &lt;code&gt;git pull&lt;/code&gt; that silently moves you onto a new vLLM release is a release-upgrade decision, not a no-op sync. Pin the recipe, review what changed, then merge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never hand-edit state files an onboarding tool owns&lt;/strong&gt; — &lt;code&gt;sandboxes.json&lt;/code&gt; and friends. Redo the supported command with the right flag instead.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat agent-suggested flags as hypotheses, not commands.&lt;/strong&gt; &lt;code&gt;--vllm-v1-disable&lt;/code&gt; was pure fiction.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Screenshot your terminal during an incident.&lt;/strong&gt; Half this post was reconstructed from them — the crash trace and &lt;code&gt;git pathspec&lt;/code&gt; error would otherwise be memory.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;./flash doctor&lt;/code&gt; + &lt;code&gt;./flash test&lt;/code&gt; after every infra event&lt;/strong&gt; (reboot, network switch, image update). The nasty failures on this box are the ones that leave a health endpoint green and the quality quietly wrong.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The Spark is a weird, wonderful machine — a desktop with server-class unified memory. It punishes you the moment you treat its memory as infinite. But the fix was never a flag: it was knowing which 48 GiB you can stop paying for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three more things the screenshots taught me
&lt;/h2&gt;

&lt;p&gt;Since publishing the first draft I OCR'd my incident screenshots — more terminal history than&lt;br&gt;
memory — and found three extra lessons:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The container port map matters.&lt;/strong&gt; &lt;code&gt;docker ps&lt;/code&gt; showed the engine publishing
&lt;code&gt;0.0.0.0:18300-&amp;gt;8000&lt;/code&gt; — vLLM listens on 8000 &lt;em&gt;inside&lt;/em&gt;, the host maps 18300. A stale
sandbox config probing &lt;code&gt;host:8000&lt;/code&gt; fails with &lt;code&gt;curl (7)&lt;/code&gt; and looks exactly like "the
engine is down" when it is up and healthy. Grep every layer's config for the old port.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;FP8 KV cache is a quality trade, not a free win.&lt;/strong&gt; One recipe's memory-budget report
warned that fp8 KV gives ~1.7x more KV tokens (1M context looks reachable) but dropped a
long-reasoning benchmark from 6/6 to 2/6 — with sparse attention, quantized keys perturb
which blocks the indexer selects. Re-validate quality on your own workload before trading
it for context length.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your host kernel may be mis-tuned for the NVIDIA driver.&lt;/strong&gt; The same report flagged
&lt;code&gt;vm.min_free_kbytes=45166&lt;/code&gt; and &lt;code&gt;vm.watermark_scale_factor=10&lt;/code&gt; at defaults — no free-page
reserve for the GPU driver. The recipe won't apply it for you (&lt;code&gt;sudo sysctl -p&lt;/code&gt; is a
one-time host step). It's recipe-agnostic: check it once, keep it forever.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;em&gt;Engine: patched vLLM v0.30.0 (&lt;a href="https://github.com/blazux/qwen3.8-Flash-DGX" rel="noopener noreferrer"&gt;blazux/qwen3.8-Flash-DGX&lt;/a&gt;, Apache-2.0) on a single DGX Spark. Checkpoint: nvidia/Qwen3.8-Flash-Next-NVFP4.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devops</category>
      <category>machinelearning</category>
      <category>docker</category>
      <category>tutorial</category>
    </item>
  </channel>
</rss>
