<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: pureGavin</title>
    <description>The latest articles on DEV Community by pureGavin (@puregavin).</description>
    <link>https://dev.to/puregavin</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4122333%2Fb09fd859-8522-4c2f-bc67-1aa1a51cebbc.png</url>
      <title>DEV Community: pureGavin</title>
      <link>https://dev.to/puregavin</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/puregavin"/>
    <language>en</language>
    <item>
      <title>Lenovo ThinkStation PGX Repair and Multi-GPU Interconnection</title>
      <dc:creator>pureGavin</dc:creator>
      <pubDate>Sat, 12 Sep 2026 16:05:16 +0000</pubDate>
      <link>https://dev.to/puregavin/lenovo-thinkstation-pgx-repair-and-multi-gpu-interconnection-3i1</link>
      <guid>https://dev.to/puregavin/lenovo-thinkstation-pgx-repair-and-multi-gpu-interconnection-3i1</guid>
      <description>&lt;h1&gt;
  
  
  Preface
&lt;/h1&gt;

&lt;p&gt;Recently I got a few PGX units for research. This is an ARM-architecture personal GPU workstation from the NVIDIA–Lenovo collaboration. In actual use it always runs into all kinds of problems. The main content of this article is repairing a recently obtained PGX whose operating system could not boot, reinstalling the OS, and then using the QSFP ports that came with the purchased machines for multi-GPU interconnection.&lt;/p&gt;

&lt;h1&gt;
  
  
  Main Text
&lt;/h1&gt;

&lt;h2&gt;
  
  
  Reinstalling the OS
&lt;/h2&gt;

&lt;p&gt;This is mentioned in the &lt;a href="https://support.lenovo.com/tw/zh/solutions/ht518087-how-to-install-dgx-spark-os-pgx-workstation" rel="noopener noreferrer"&gt;Lenovo official documentation&lt;/a&gt;, but the most important thing is that it does not provide the ISO image for reinstallation. Many people may download the image directly from the NVIDIA official website. I looked into this: the GB10 machine image provided by NVIDIA is not the same as the Lenovo-customized image. Here I provide a &lt;a href="https://download.lenovo.com/km/media/attachment/DGXOS_7_4_0_GA2_0_PGX.iso" rel="noopener noreferrer"&gt;download link&lt;/a&gt; for the PGX image, but if your local network has issues that prevent access to the download link above, you can also try downloading via the magnet link(magnet:?xt=urn:btih:3eec242622ead59b31977c59824c83fc73bd9a12&amp;amp;dn=DGXOS_7_4_0_GA2_0_PGX.iso&amp;amp;tr=udp%3A%2F%2Ftracker.opentrackr.org%3A1337%2Fannounce) I uploaded to my own server (please weigh the security of the image file yourself; I cannot make any guarantees, and any subsequent security issues are not my responsibility)&lt;/p&gt;

&lt;p&gt;Reinstalling the OS itself is not difficult, and the Lenovo official link given above is also quite detailed. Here, to prevent some people from being unable to access it due to network issues, I will briefly go over it. Assume you have already downloaded the PGX ISO image file, then download the rufus tool from rufus's &lt;a href="https://github.com/pbatard/rufus" rel="noopener noreferrer"&gt;GitHub official repository&lt;/a&gt;, then insert a USB drive (capacity of at least 16G), open the rufus tool, select the USB drive, then select the downloaded ISO image, leave everything else unchanged at the defaults, then click Start directly&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.com%2FpureGavin%2Fphoto%2Fblob%2Fmain%2FNAS%2Flenovo%2520thinkstation%2520PGX%25E4%25BF%25AE%25E5%25A4%258D%25E4%25B8%258E%25E5%25A4%259A%25E5%258D%25A1%25E8%25BF%259E%25E6%258E%25A5%2F1.png%3Fraw%3Dtrue" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.com%2FpureGavin%2Fphoto%2Fblob%2Fmain%2FNAS%2Flenovo%2520thinkstation%2520PGX%25E4%25BF%25AE%25E5%25A4%258D%25E4%25B8%258E%25E5%25A4%259A%25E5%258D%25A1%25E8%25BF%259E%25E6%258E%25A5%2F1.png%3Fraw%3Dtrue" alt="1" width="590" height="986"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After some time, the USB drive becomes a boot drive, then just click Close&lt;/p&gt;

&lt;p&gt;Neither the PGX's external Type-C ports nor the included USB adapter ports work with 2.4G wireless keyboards; you have to use an old-fashioned USB-wired keyboard. After powering on, keep pressing the Delete key to enter the boot page, press Tab on the keyboard, select the bss option, enter the hardware boot order, and select USB drive boot&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.com%2FpureGavin%2Fphoto%2Fblob%2Fmain%2FNAS%2Flenovo%2520thinkstation%2520PGX%25E4%25BF%25AE%25E5%25A4%258D%25E4%25B8%258E%25E5%25A4%259A%25E5%258D%25A1%25E8%25BF%259E%25E6%258E%25A5%2F2.png%3Fraw%3Dtrue" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.com%2FpureGavin%2Fphoto%2Fblob%2Fmain%2FNAS%2Flenovo%2520thinkstation%2520PGX%25E4%25BF%25AE%25E5%25A4%258D%25E4%25B8%258E%25E5%25A4%259A%25E5%258D%25A1%25E8%25BF%259E%25E6%258E%25A5%2F2.png%3Fraw%3Dtrue" alt="2" width="537" height="368"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After entering the GRUB menu, select the DGX install option&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.com%2FpureGavin%2Fphoto%2Fblob%2Fmain%2FNAS%2Flenovo%2520thinkstation%2520PGX%25E4%25BF%25AE%25E5%25A4%258D%25E4%25B8%258E%25E5%25A4%259A%25E5%258D%25A1%25E8%25BF%259E%25E6%258E%25A5%2F3.jpg%3Fraw%3Dtrue" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.com%2FpureGavin%2Fphoto%2Fblob%2Fmain%2FNAS%2Flenovo%2520thinkstation%2520PGX%25E4%25BF%25AE%25E5%25A4%258D%25E4%25B8%258E%25E5%25A4%259A%25E5%258D%25A1%25E8%25BF%259E%25E6%258E%25A5%2F3.jpg%3Fraw%3Dtrue" alt="3" width="1269" height="816"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then click to install the DGX system&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.com%2FpureGavin%2Fphoto%2Fblob%2Fmain%2FNAS%2Flenovo%2520thinkstation%2520PGX%25E4%25BF%25AE%25E5%25A4%258D%25E4%25B8%258E%25E5%25A4%259A%25E5%258D%25A1%25E8%25BF%259E%25E6%258E%25A5%2F4.png%3Fraw%3Dtrue" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fgithub.com%2FpureGavin%2Fphoto%2Fblob%2Fmain%2FNAS%2Flenovo%2520thinkstation%2520PGX%25E4%25BF%25AE%25E5%25A4%258D%25E4%25B8%258E%25E5%25A4%259A%25E5%258D%25A1%25E8%25BF%259E%25E6%258E%25A5%2F4.png%3Fraw%3Dtrue" alt="4" width="692" height="408"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then the operating system will install automatically; you just need to wait. After installation completes, you will enter the familiar initialization interface to configure language, timezone, account password, and so on&lt;/p&gt;

&lt;h2&gt;
  
  
  Multi-machine interconnection
&lt;/h2&gt;

&lt;p&gt;Although it is titled multi-machine interconnection, in fact I only have two machines here, because the ports on the back of GB10-series machines use QSFP (200G) interfaces, and switches for this interface are hard for individuals to buy or are very expensive, so I only demonstrate two machines here. However, from the configuration principle, connecting more than two should be supported&lt;/p&gt;

&lt;p&gt;The entire tutorial can also be found on the &lt;a href="https://developer.nvidia.cn/build-spark/connect-two-sparks#i7njvai" rel="noopener noreferrer"&gt;NVIDIA official website&lt;/a&gt;, but as above, because some people may have network issues, I will also recap it here&lt;/p&gt;

&lt;p&gt;The first step is to confirm whether the usernames on the two machines are the same; just use the &lt;code&gt;whoami&lt;/code&gt; command. If the usernames on the two machines are different, then you need to create two users with the same username in the system. Basic Linux commands will not be shown here; you can search for them yourselves&lt;/p&gt;

&lt;p&gt;The second step is also very simple: use the QSFP ports to connect the two machines. Be careful not to use force when connecting; the ports are keyed, and you cannot insert them the wrong way. After connecting, run the &lt;code&gt;ibdev2netdev&lt;/code&gt; command, and you will see the following result&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;lenovo@thinkstationpgx-3262:~/PGX&lt;span class="nv"&gt;$ &lt;/span&gt;ibdev2netdev
rocep1s0f0 port 1 &lt;span class="o"&gt;==&amp;gt;&lt;/span&gt; enp1s0f0np0 &lt;span class="o"&gt;(&lt;/span&gt;Up&lt;span class="o"&gt;)&lt;/span&gt;
rocep1s0f1 port 1 &lt;span class="o"&gt;==&amp;gt;&lt;/span&gt; enp1s0f1np1 &lt;span class="o"&gt;(&lt;/span&gt;Down&lt;span class="o"&gt;)&lt;/span&gt;
roceP2p1s0f0 port 1 &lt;span class="o"&gt;==&amp;gt;&lt;/span&gt; enP2p1s0f0np0 &lt;span class="o"&gt;(&lt;/span&gt;Up&lt;span class="o"&gt;)&lt;/span&gt;
roceP2p1s0f1 port 1 &lt;span class="o"&gt;==&amp;gt;&lt;/span&gt; enP2p1s0f1np1 &lt;span class="o"&gt;(&lt;/span&gt;Down&lt;span class="o"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The third step is configuring the machines' &lt;code&gt;netplan&lt;/code&gt;. The official documentation provides three configuration methods (one automatic and two manual); here I used automatic configuration&lt;/p&gt;

&lt;p&gt;First write a yaml configuration file under the &lt;code&gt;netplan&lt;/code&gt; directory&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo tee&lt;/span&gt; /etc/netplan/40-cx7.yaml &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /dev/null &lt;span class="o"&gt;&amp;lt;&amp;lt;&lt;/span&gt;&lt;span class="no"&gt;EOF&lt;/span&gt;&lt;span class="sh"&gt;
network:
  version: 2
  ethernets:
    enp1s0f0np0:
      link-local: [ ipv4 ]
    enp1s0f1np1:
      link-local: [ ipv4 ]
&lt;/span&gt;&lt;span class="no"&gt;EOF
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then configure the corresponding permissions and deploy&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Configure permissions&lt;/span&gt;
&lt;span class="nb"&gt;sudo chmod &lt;/span&gt;600 /etc/netplan/40-cx7.yaml

&lt;span class="c"&gt;# Apply configuration&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;netplan apply
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After this step is complete, using the &lt;code&gt;ifconfig&lt;/code&gt; command or the &lt;code&gt;ip addr&lt;/code&gt; command you will see that the corresponding NICs from step two have been automatically assigned IP addresses; note that this step needs to be done on both machines&lt;/p&gt;

&lt;p&gt;The fourth step is configuring passwordless SSH certificates. NVIDIA official provides a &lt;a href="https://github.com/NVIDIA/dgx-spark-playbooks/blob/main/nvidia/connect-two-sparks/assets/discover-sparks" rel="noopener noreferrer"&gt;script&lt;/a&gt;, with the following content:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#&lt;/span&gt;
&lt;span class="c"&gt;# SPDX-FileCopyrightText: Copyright (c) 1993-2025 NVIDIA CORPORATION &amp;amp; AFFILIATES. All rights reserved.&lt;/span&gt;
&lt;span class="c"&gt;# SPDX-License-Identifier: Apache-2.0&lt;/span&gt;
&lt;span class="c"&gt;#&lt;/span&gt;
&lt;span class="c"&gt;# Licensed under the Apache License, Version 2.0 (the "License");&lt;/span&gt;
&lt;span class="c"&gt;# you may not use this file except in compliance with the License.&lt;/span&gt;
&lt;span class="c"&gt;# You may obtain a copy of the License at&lt;/span&gt;
&lt;span class="c"&gt;#&lt;/span&gt;
&lt;span class="c"&gt;# http://www.apache.org/licenses/LICENSE-2.0&lt;/span&gt;
&lt;span class="c"&gt;#&lt;/span&gt;
&lt;span class="c"&gt;# Unless required by applicable law or agreed to in writing, software&lt;/span&gt;
&lt;span class="c"&gt;# distributed under the License is distributed on an "AS IS" BASIS,&lt;/span&gt;
&lt;span class="c"&gt;# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.&lt;/span&gt;
&lt;span class="c"&gt;# See the License for the specific language governing permissions and&lt;/span&gt;
&lt;span class="c"&gt;# limitations under the License.&lt;/span&gt;
&lt;span class="c"&gt;#&lt;/span&gt;
&lt;span class="c"&gt;#!/bin/env bash&lt;/span&gt;

&lt;span class="c"&gt;# discover-sparks&lt;/span&gt;
&lt;span class="c"&gt;# Discover available systems using avahi-browse and generate MPI hosts file&lt;/span&gt;
&lt;span class="c"&gt;# Searches all active interfaces automatically&lt;/span&gt;
&lt;span class="c"&gt;#&lt;/span&gt;
&lt;span class="c"&gt;# Usage: bash ./discover-sparks&lt;/span&gt;

&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-euo&lt;/span&gt; pipefail

&lt;span class="c"&gt;# Check if running as root or with sudo&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="nv"&gt;$EUID&lt;/span&gt; &lt;span class="nt"&gt;-eq&lt;/span&gt; 0 &lt;span class="o"&gt;]]&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;SUDO_USER&lt;/span&gt;&lt;span class="k"&gt;:-}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Error: This script should not be run as root or with sudo"&lt;/span&gt;
    &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Please run as a regular user"&lt;/span&gt;
    &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="c"&gt;# Dynamically get interface names from ibdev2netdev output&lt;/span&gt;
&lt;span class="c"&gt;# Use ibdev2netdev to list Infiniband devices and their network interfaces.&lt;/span&gt;
&lt;span class="c"&gt;# The awk command searches for lines containing 'Up)' (i.e., interfaces that are up)&lt;/span&gt;
&lt;span class="c"&gt;# and prints the 5th field, which is the interface name (e.g., enp1s0f0np0).&lt;/span&gt;
&lt;span class="c"&gt;# The tr command removes any parentheses from the output.&lt;/span&gt;
&lt;span class="nv"&gt;INTERFACES&lt;/span&gt;&lt;span class="o"&gt;=(&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;ibdev2netdev | &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'/Up\)/ {print $5}'&lt;/span&gt; | &lt;span class="nb"&gt;tr&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'()'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="o"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="k"&gt;${#&lt;/span&gt;&lt;span class="nv"&gt;INTERFACES&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt; &lt;span class="nt"&gt;-eq&lt;/span&gt; 0 &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"ERROR: No active interfaces found via ibdev2netdev."&lt;/span&gt;
    &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="c"&gt;# Create temporary file for processing&lt;/span&gt;
&lt;span class="nv"&gt;TEMP_FILE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;mktemp&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="nb"&gt;trap&lt;/span&gt; &lt;span class="s1"&gt;'rm -f "$TEMP_FILE"'&lt;/span&gt; EXIT

&lt;span class="c"&gt;# Check if avahi-browse is available&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="nb"&gt;command&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt; avahi-browse &amp;amp;&amp;gt; /dev/null&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Error: avahi-browse not found. Please install avahi-utils package."&lt;/span&gt;
    &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="c"&gt;# Run avahi-browse and filter for SSH services on specified interfaces&lt;/span&gt;
&lt;span class="c"&gt;# -p: parseable output&lt;/span&gt;
&lt;span class="c"&gt;# -r: resolve host names and addresses&lt;/span&gt;
&lt;span class="c"&gt;# -f: terminate after dumping all entries available at startup&lt;/span&gt;
&lt;span class="nv"&gt;avahi_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;avahi-browse &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="nt"&gt;-t&lt;/span&gt; _ssh._tcp 2&amp;gt;/dev/null&lt;span class="si"&gt;)&lt;/span&gt;

&lt;span class="c"&gt;# Filter for both interfaces&lt;/span&gt;
&lt;span class="nv"&gt;found_services&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;false
&lt;/span&gt;&lt;span class="k"&gt;for &lt;/span&gt;interface &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;INTERFACES&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
    if &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$avahi_output&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$interface&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TEMP_FILE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
        &lt;/span&gt;&lt;span class="nv"&gt;found_services&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;true
    &lt;/span&gt;&lt;span class="k"&gt;fi
done

if&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$found_services&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;false&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Warning: No services found on any specified interface"&lt;/span&gt;
    &lt;span class="nb"&gt;exit &lt;/span&gt;0
&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="c"&gt;# Extract IPv4 addresses from the avahi-browse output&lt;/span&gt;
&lt;span class="c"&gt;# Format: =;interface;IPv4;hostname\032service;description;local;fqdn;ip_address;port;&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="s2"&gt;"^="&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TEMP_FILE&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="s2"&gt;"IPv4"&lt;/span&gt; | &lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="nv"&gt;IFS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;';'&lt;/span&gt; &lt;span class="nb"&gt;read&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; prefix interface protocol hostname_service description &lt;span class="nb"&gt;local &lt;/span&gt;fqdn ip_address port rest&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
    &lt;span class="c"&gt;# Clean up any trailing data&lt;/span&gt;
    &lt;span class="nv"&gt;clean_ip&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$ip_address&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | &lt;span class="nb"&gt;sed&lt;/span&gt; &lt;span class="s1"&gt;'s/;.*$//'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;

    &lt;span class="c"&gt;# Validate IP address format&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="nv"&gt;$clean_ip&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;~ ^[0-9]&lt;span class="o"&gt;{&lt;/span&gt;1,3&lt;span class="o"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\.&lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;0-9]&lt;span class="o"&gt;{&lt;/span&gt;1,3&lt;span class="o"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\.&lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;0-9]&lt;span class="o"&gt;{&lt;/span&gt;1,3&lt;span class="o"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\.&lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;0-9]&lt;span class="o"&gt;{&lt;/span&gt;1,3&lt;span class="o"&gt;}&lt;/span&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
        &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$clean_ip&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TEMP_FILE&lt;/span&gt;&lt;span class="s2"&gt;.sorted"&lt;/span&gt;
        &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Found: &lt;/span&gt;&lt;span class="nv"&gt;$clean_ip&lt;/span&gt;&lt;span class="s2"&gt; (&lt;/span&gt;&lt;span class="nv"&gt;$fqdn&lt;/span&gt;&lt;span class="s2"&gt;)"&lt;/span&gt;
    &lt;span class="k"&gt;else
        &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Warning: Invalid IP format: &lt;/span&gt;&lt;span class="nv"&gt;$clean_ip&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
    &lt;span class="k"&gt;fi
done&lt;/span&gt;

&lt;span class="c"&gt;# Sort and remove duplicates&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TEMP_FILE&lt;/span&gt;&lt;span class="s2"&gt;.sorted"&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TEMP_FILE&lt;/span&gt;&lt;span class="s2"&gt;.sorted"&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TEMP_FILE&lt;/span&gt;&lt;span class="s2"&gt;.sorted"&lt;/span&gt;
&lt;span class="k"&gt;else
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"No IPv4 addresses found."&lt;/span&gt;
    &lt;span class="nb"&gt;exit &lt;/span&gt;1
&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="c"&gt;# Generate a shared SSH key if it doesn't exist&lt;/span&gt;
&lt;span class="nv"&gt;SHARED_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/.ssh/id_ed25519_shared"&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SHARED_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Generating shared SSH key for all nodes..."&lt;/span&gt;
    ssh-keygen &lt;span class="nt"&gt;-t&lt;/span&gt; ed25519 &lt;span class="nt"&gt;-N&lt;/span&gt; &lt;span class="s2"&gt;""&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SHARED_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="nt"&gt;-C&lt;/span&gt; &lt;span class="s2"&gt;"shared-cluster-key"&lt;/span&gt;
&lt;span class="k"&gt;fi

&lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;""&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Setting up shared SSH access across all nodes..."&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"You may be prompted for your password on each node."&lt;/span&gt;

&lt;span class="c"&gt;# Ensure local .ssh directory exists with correct permissions&lt;/span&gt;
&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/.ssh"&lt;/span&gt;
&lt;span class="nb"&gt;chmod &lt;/span&gt;700 &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/.ssh"&lt;/span&gt;

&lt;span class="c"&gt;# Add shared public key to local authorized_keys&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-qF&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SHARED_KEY&lt;/span&gt;&lt;span class="s2"&gt;.pub"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/.ssh/authorized_keys"&lt;/span&gt; 2&amp;gt;/dev/null&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;cat&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SHARED_KEY&lt;/span&gt;&lt;span class="s2"&gt;.pub"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/.ssh/authorized_keys"&lt;/span&gt;
    &lt;span class="nb"&gt;chmod &lt;/span&gt;600 &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/.ssh/authorized_keys"&lt;/span&gt;
    &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"  ✓ Added shared public key to local authorized_keys"&lt;/span&gt;
&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="c"&gt;# Distribute shared key to all remote nodes&lt;/span&gt;
&lt;span class="k"&gt;while &lt;/span&gt;&lt;span class="nb"&gt;read&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; node_ip&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
    if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$node_ip&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
        &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Configuring &lt;/span&gt;&lt;span class="nv"&gt;$node_ip&lt;/span&gt;&lt;span class="s2"&gt;..."&lt;/span&gt;

        &lt;span class="c"&gt;# Copy shared key to remote node and set up authorized_keys&lt;/span&gt;
        &lt;span class="k"&gt;if &lt;/span&gt;scp &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;StrictHostKeyChecking&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;accept-new &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SHARED_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SHARED_KEY&lt;/span&gt;&lt;span class="s2"&gt;.pub"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$USER&lt;/span&gt;&lt;span class="s2"&gt;@&lt;/span&gt;&lt;span class="nv"&gt;$node_ip&lt;/span&gt;&lt;span class="s2"&gt;:~/.ssh/"&lt;/span&gt; &amp;amp;&amp;gt;/dev/null&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
            &lt;/span&gt;ssh &lt;span class="nt"&gt;-n&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;StrictHostKeyChecking&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;accept-new &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$USER&lt;/span&gt;&lt;span class="s2"&gt;@&lt;/span&gt;&lt;span class="nv"&gt;$node_ip&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"
                chmod 700 ~/.ssh
                chmod 600 ~/.ssh/id_ed25519_shared
                chmod 644 ~/.ssh/id_ed25519_shared.pub

                # Add shared public key to authorized_keys if not present
                if ! grep -qF &lt;/span&gt;&lt;span class="se"&gt;\"\$&lt;/span&gt;&lt;span class="s2"&gt;(cat ~/.ssh/id_ed25519_shared.pub)&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt; ~/.ssh/authorized_keys 2&amp;gt;/dev/null; then
                    cat ~/.ssh/id_ed25519_shared.pub &amp;gt;&amp;gt; ~/.ssh/authorized_keys
                    chmod 600 ~/.ssh/authorized_keys
                fi

                # Create/update SSH config to use shared key by default
                if ! grep -q 'IdentityFile.*id_ed25519_shared' ~/.ssh/config 2&amp;gt;/dev/null; then
                    echo 'Host *' &amp;gt;&amp;gt; ~/.ssh/config
                    echo '    IdentityFile ~/.ssh/id_ed25519_shared' &amp;gt;&amp;gt; ~/.ssh/config
                    chmod 600 ~/.ssh/config
                fi
            "&lt;/span&gt; &amp;amp;&amp;gt;/dev/null

            &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"  ✓ Successfully configured &lt;/span&gt;&lt;span class="nv"&gt;$node_ip&lt;/span&gt;&lt;span class="s2"&gt; with shared key"&lt;/span&gt;
        &lt;span class="k"&gt;else
            &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"  ✗ Failed to configure &lt;/span&gt;&lt;span class="nv"&gt;$node_ip&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
        &lt;span class="k"&gt;fi
    fi
done&lt;/span&gt; &amp;lt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TEMP_FILE&lt;/span&gt;&lt;span class="s2"&gt;.sorted"&lt;/span&gt;

&lt;span class="c"&gt;# Update local SSH config to use shared key&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt; &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="s1"&gt;'IdentityFile.*id_ed25519_shared'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/.ssh/config"&lt;/span&gt; 2&amp;gt;/dev/null&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;touch&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/.ssh/config"&lt;/span&gt;
    &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'Host *'&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/.ssh/config"&lt;/span&gt;
    &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'    IdentityFile ~/.ssh/id_ed25519_shared'&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/.ssh/config"&lt;/span&gt;
    &lt;span class="nb"&gt;chmod &lt;/span&gt;600 &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/.ssh/config"&lt;/span&gt;
    &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"  ✓ Updated local SSH config to use shared key"&lt;/span&gt;
&lt;span class="k"&gt;fi

&lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;""&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Shared SSH setup complete!"&lt;/span&gt;
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"All nodes can now SSH to each other using the shared key (id_ed25519_shared)."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here if you encounter this problem&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;Error: This script should not be run as root or with &lt;span class="nb"&gt;sudo
&lt;/span&gt;Please run as a regular user
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You need to start a new terminal, then without doing anything else, run this script directly&lt;/p&gt;

&lt;p&gt;After this script executes successfully, you can use the following command to test the result; the IP address is the output from step three&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# The return result is the hostname of the machine at the corresponding IP address&lt;/span&gt;
ssh &amp;lt;ip address&amp;gt; &lt;span class="nb"&gt;hostname&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At this point, the steps for connecting two PGX units via QSFP are complete; next is the software-level work&lt;/p&gt;

&lt;h2&gt;
  
  
  Building a multi-GPU cluster
&lt;/h2&gt;

&lt;p&gt;This involves some specialized terms (Ray cluster, NCCL communication), but the purpose of this article is not terminology explanation, but how to use these things directly (even so, some simple principle introduction is still needed); the large-model environment in this article is docker+vllm&lt;/p&gt;

&lt;p&gt;Here we use the latest &lt;code&gt;Qwen/Qwen3.8-Flash-Next-FP8&lt;/code&gt;, mainly because BF16 is &lt;strong&gt;not viable&lt;/strong&gt; on these two machines—not an ordinary OOM, but dragging the local machine to death and reboot; what finally went into production is FP8, with a topology of TP=2 across two single-GPU nodes, the API on the local machine at &lt;code&gt;http://localhost:20001&lt;/code&gt;, and the external model name &lt;code&gt;qwen3.8-flash-next&lt;/code&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Why two machines are required
&lt;/h3&gt;

&lt;p&gt;GB10 is a unified memory architecture. The "GPU memory" reported by &lt;code&gt;nvidia-smi&lt;/code&gt; and the physical memory in &lt;code&gt;/proc/meminfo&lt;/code&gt; are the same pool; each machine physically has only &lt;strong&gt;119 GiB&lt;/strong&gt;. The so-called "requesting GPU memory" takes the kernel path through &lt;code&gt;nv_alloc_system_pages&lt;/code&gt;, that is, asking the system for pages.&lt;/p&gt;

&lt;p&gt;The two checkpoints of Qwen3.8-Flash-Next:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Checkpoint&lt;/th&gt;
&lt;th&gt;Weight size (as read by the loader)&lt;/th&gt;
&lt;th&gt;Disk occupancy&lt;/th&gt;
&lt;th&gt;TP=2 per-node weights&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;Qwen/Qwen3.8-Flash-Next&lt;/code&gt; (BF16)&lt;/td&gt;
&lt;td&gt;335.28 GiB&lt;/td&gt;
&lt;td&gt;336G&lt;/td&gt;
&lt;td&gt;167.6 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Qwen/Qwen3.8-Flash-Next-FP8&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;172.78 GiB&lt;/td&gt;
&lt;td&gt;173G&lt;/td&gt;
&lt;td&gt;86.4 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;BF16 is 167.6 GiB per node, &lt;strong&gt;48.6 GiB&lt;/strong&gt; more than 119 GiB. The gap cannot be made up by paging; the driver will hold a lock in the kernel and repeatedly reclaim and retry, and the whole machine dies along with it. This path has already been walked in actual testing (mentioned later).&lt;/p&gt;

&lt;p&gt;FP8 weights per node actually occupy &lt;strong&gt;86.33 GiB&lt;/strong&gt;. KV and GDN state still need to be added. In &lt;code&gt;config.json&lt;/code&gt;, &lt;code&gt;num_hidden_layers=48&lt;/code&gt;, &lt;code&gt;full_attention_interval=4&lt;/code&gt;, that is, 12 QSA full-attention layers and 36 Gated DeltaNet layers; &lt;code&gt;num_key_value_heads=2&lt;/code&gt;, &lt;code&gt;head_dim=256&lt;/code&gt;. KV counted as BF16:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;12 layers × 2 KV heads × 256 dim × 2 (K/V) × 2 bytes = 24 KiB / token
After TP=2, 12 KiB / token per node
262144 context, a single full-length sequence ≈ 3 GiB / node
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;GDN recurrent state: 36 layers × 48 V heads × 128 × 128 × 2 bytes = 54 MiB / sequence, 27 MiB after TP=2; &lt;code&gt;--max-num-seqs 8&lt;/code&gt; totals 216 MiB, which can be ignored. Weights 86.4 GiB plus KV and a bit of host overhead come to about 95 GiB, which fits in 119 GiB, so &lt;code&gt;--max-model-len 262144&lt;/code&gt; can be kept.&lt;/p&gt;

&lt;p&gt;In the final CUDA graph round, the local KV cache was &lt;strong&gt;7.62 GiB / 541,401 tokens&lt;/strong&gt;, maximum concurrency for 262144 context &lt;strong&gt;2.07x&lt;/strong&gt;; the peer KV was &lt;strong&gt;6.78 GiB&lt;/strong&gt;. The two sides differ slightly because the idle memory on the two machines was not exactly the same at the time, and vLLM each took a cut according to &lt;code&gt;--gpu-memory-utilization 0.80&lt;/code&gt; (budget 95.7 GiB)&lt;/p&gt;

&lt;h3&gt;
  
  
  There is only one topology choice
&lt;/h3&gt;

&lt;p&gt;This architecture has N-gram embedding (&lt;code&gt;ngram_vocab_size_base=20000000&lt;/code&gt;, attached at layer 2). vLLM &lt;strong&gt;does not support pipeline parallelism&lt;/strong&gt; for it. PLE offloading the n-gram table to CPU might be tryable on a single machine, but &lt;strong&gt;distributed mode is not supported&lt;/strong&gt;. Both machines are single-GPU, so only one topology remains: split into two shares with tensor parallelism, one GPU per machine.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
  Client["Client :20001"] --&amp;gt; Rank0
  subgraph rank0node ["thinkstationpgx-30 / rank 0"]
    Rank0["APIServer + EngineCore / TP rank 0"]
  end
  subgraph rank1node ["thinkstationpgx-32 / rank 1"]
    Rank1["headless Worker / TP rank 1"]
  end
  Rank0 --&amp;gt;|"RoCE NCCL + Gloo"| Rank1&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The local machine is the master: &lt;code&gt;169.254.94.252&lt;/code&gt;, &lt;code&gt;NODE_RANK=0&lt;/code&gt;, running the full APIServer and EngineCore, with HTTP bound to &lt;code&gt;0.0.0.0:20001&lt;/code&gt;. The peer is the follower: &lt;code&gt;169.254.159.206&lt;/code&gt;, &lt;code&gt;NODE_RANK=1&lt;/code&gt;, running only a &lt;code&gt;--headless&lt;/code&gt; Worker, not serving HTTP externally. NCCL goes over the two RoCE logical NICs on QSFP, and the Gloo / TCP control plane goes over &lt;code&gt;enp1s0f0np0&lt;/code&gt;. The default route remains on &lt;code&gt;enP7s7&lt;/code&gt; (the &lt;code&gt;192.168.40.x&lt;/code&gt; LAN segment); &lt;strong&gt;do not change the default gateway to QSFP&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The two machines use the same compose, distinguishing identity via &lt;code&gt;NODE_RANK&lt;/code&gt; and &lt;code&gt;VLLM_HOST_IP&lt;/code&gt; in each machine's &lt;code&gt;.env&lt;/code&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Three checks before starting work
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. Kernel-mode and user-mode driver versions must match.&lt;/strong&gt; The peer once had kernel &lt;code&gt;580.159.03&lt;/code&gt; and user-mode &lt;code&gt;580.173.02&lt;/code&gt;; the machine had already been running continuously for fifty-plus days, and GPU initialization failed inside Docker. It only recovered after rebooting the peer. Both the local machine and the peer must be able to see the GPU inside the container; do not just glance at &lt;code&gt;nvidia-smi&lt;/code&gt; on the host.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Firewall must allow the QSFP network segments and the API port.&lt;/strong&gt; The local UFW defaults to DROP, which will block &lt;code&gt;29501/tcp&lt;/code&gt; used by torch distributed. You need:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;ufw allow from 169.254.0.0/16
&lt;span class="nb"&gt;sudo &lt;/span&gt;ufw allow from 192.168.101.0/24
&lt;span class="nb"&gt;sudo &lt;/span&gt;ufw allow 20001/tcp
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;UFW was not enabled on the peer at the time, so no change was needed. The rules must survive reboot.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Both RoCE logical NICs must be up.&lt;/strong&gt; The physical port is QSFP port 0, split into two logical NICs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Ethernet name&lt;/th&gt;
&lt;th&gt;RoCE name&lt;/th&gt;
&lt;th&gt;Local&lt;/th&gt;
&lt;th&gt;Peer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;enp1s0f0np0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;rocep1s0f0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;169.254.94.252/16&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;169.254.159.206/16&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;enP2p1s0f0np0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;roceP2p1s0f0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;192.168.101.10/24&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;192.168.101.11/24&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;GID index &lt;strong&gt;3&lt;/strong&gt; is RoCEv2 + IPv4. &lt;code&gt;ib_write_bw&lt;/code&gt; reached &lt;strong&gt;109 Gb/s&lt;/strong&gt; on each of the two. NCCL uses &lt;code&gt;NCCL_IB_HCA==rocep1s0f0,roceP2p1s0f0&lt;/code&gt; (the equals sign is NCCL's exact-match syntax) plus &lt;code&gt;NCCL_IB_MERGE_NICS=1&lt;/code&gt;, merging the two into a single 200 GbE-class transport&lt;/p&gt;

&lt;h3&gt;
  
  
  Image and weights
&lt;/h3&gt;

&lt;p&gt;Qwen3.8-Flash-Next must use the dedicated image &lt;code&gt;vllm/vllm-openai:qwen38-flash-next&lt;/code&gt; (arm64). The generic &lt;code&gt;v0.28.0&lt;/code&gt; &lt;strong&gt;does not recognize&lt;/strong&gt; this architecture; pulling the wrong image will exit directly at the model initialization stage. The version string that actually ran in this round is &lt;code&gt;v0.1.dev20073+g8e685d198&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;One of the GB10s here has very slow network, so both the image and the weights were prepared on the local machine and then pushed over via QSFP. The working directory on both sides is &lt;code&gt;/home/lenovo/vllm&lt;/code&gt;, and the Hugging Face cache is mounted at &lt;code&gt;/home/lenovo/vllm/huggingface&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;For downloading you can run &lt;code&gt;hf download&lt;/code&gt; with a generic image already on the local machine, just bind the cache directory to the path above; there is no need to download inside the dedicated image. The HF cache owner is root, and &lt;code&gt;trees/*.json&lt;/code&gt; is often mode &lt;code&gt;600&lt;/code&gt;, which the local user cannot read, and rsync will fail on these small files. Before transferring, first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo chmod&lt;/span&gt; &lt;span class="nt"&gt;-R&lt;/span&gt; a+rX /home/lenovo/vllm/huggingface/hub/models--Qwen--Qwen3.8-Flash-Next-FP8
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three notes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do not go via hostname or &lt;code&gt;192.168.40.x&lt;/code&gt;; traffic must go via &lt;code&gt;169.254.159.206&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Do not use &lt;code&gt;rsync -L&lt;/code&gt;. HF's &lt;code&gt;snapshots/&lt;/code&gt; are symlinks pointing to &lt;code&gt;blobs/&lt;/code&gt;; dereferencing will double the volume.&lt;/li&gt;
&lt;li&gt;Do not enable compression for rsync / SSH. On 200 GbE, compression will only saturate the CPU and leave the bandwidth unused.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After &lt;code&gt;docker load&lt;/code&gt;, the peer's image ID may differ from the local machine's; this is because load rewrote the local image metadata. As long as the RootFS layers match, it can be used&lt;/p&gt;

&lt;h3&gt;
  
  
  Key design of the compose
&lt;/h3&gt;

&lt;p&gt;The complete file is in the appendix. Here I only cover "why it must be written this way"; each item corresponds to a failure that already happened.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Host network, InfiniBand devices, and locked-page permissions.&lt;/strong&gt; &lt;code&gt;network_mode: host&lt;/code&gt; is because NCCL needs to get the host NICs directly; a mapping like &lt;code&gt;8000:20001&lt;/code&gt; inside the container would make the NCCL handshake go through the wrong namespace. vLLM itself uses &lt;code&gt;--host 0.0.0.0 --port 20001&lt;/code&gt;, consistent with the external port of other compose files on the local machine (for example &lt;code&gt;qwen3.6-35b-compose.yml&lt;/code&gt;). Without &lt;code&gt;/dev/infiniband&lt;/code&gt;, &lt;code&gt;IPC_LOCK&lt;/code&gt;, and &lt;code&gt;memlock: -1&lt;/code&gt;, RoCE falls back to sockets and bandwidth drops to unusable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Force Gloo to IPv4, and turn off libuv.&lt;/strong&gt; The first distributed handshake hung on c10d's IPv6 socket timeout. The environment variables are:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;NCCL_SOCKET_FAMILY&lt;/span&gt;: &lt;span class="n"&gt;AF_INET&lt;/span&gt;
&lt;span class="n"&gt;GLOO_SOCKET_FAMILY&lt;/span&gt;: &lt;span class="n"&gt;AF_INET&lt;/span&gt;
&lt;span class="n"&gt;GLOO_SOCKET_IFNAME&lt;/span&gt;: &lt;span class="n"&gt;enp1s0f0np0&lt;/span&gt;
&lt;span class="n"&gt;GLOO_USE_LIBUV&lt;/span&gt;: &lt;span class="s2"&gt;"0"&lt;/span&gt;
&lt;span class="n"&gt;TORCH_GLOO_USE_LIBUV&lt;/span&gt;: &lt;span class="s2"&gt;"0"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The control plane goes over &lt;code&gt;enp1s0f0np0&lt;/code&gt;, the data plane over the two RoCE NICs. &lt;code&gt;NCCL_IB_GID_INDEX=3&lt;/code&gt; corresponds to RoCEv2+IPv4; &lt;code&gt;NCCL_CUMEM_ENABLE=0&lt;/code&gt; is the safe choice on this platform.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Container memory hard limit 105G, explicit&lt;/strong&gt; &lt;code&gt;restart: "no"&lt;/code&gt;&lt;strong&gt;.&lt;/strong&gt; Both machines are cgroup v2 + systemd driver; driver allocations are counted into the container memcg. When over limit it first reclaims page cache, then OOM-kills the &lt;strong&gt;container&lt;/strong&gt;, rather than dragging the whole machine to death. &lt;code&gt;restart: "no"&lt;/code&gt; avoids automatically bringing it back up after a crash and repeatedly slamming memory. &lt;code&gt;--gpu-memory-utilization 0.80&lt;/code&gt; gives about 95.7 GiB of GPU budget, leaving about 9 GiB between that and the 105G cap for the Python / torch host. &lt;code&gt;--distributed-timeout-seconds 3600&lt;/code&gt; is because loading 86 GiB on two nodes has a time skew (peer about 300 seconds, local about 850 seconds), and the default 600-second NCCL timeout is tight.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Wrap the entrypoint in a layer of bash; rank 1 must add&lt;/strong&gt; &lt;code&gt;--headless&lt;/code&gt;&lt;strong&gt;.&lt;/strong&gt; The image's default &lt;code&gt;ENTRYPOINT&lt;/code&gt; is &lt;code&gt;["vllm","serve"]&lt;/code&gt;. Compose needs to interpolate &lt;code&gt;NODE_RANK&lt;/code&gt; / &lt;code&gt;MODEL_ID&lt;/code&gt; from each machine's &lt;code&gt;.env&lt;/code&gt;, so it is changed to &lt;code&gt;/bin/bash -lc&lt;/code&gt;. If rank 1 runs a full EngineCore, it will blow up in &lt;code&gt;_initialize_kv_caches&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AssertionError: collective_rpc should not be called on follower node
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The follower only acts as a Worker. HTTP is only on rank 0.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Compilation flags.&lt;/strong&gt; First get it running with &lt;code&gt;--enforce-eager&lt;/code&gt; (the dual-Spark configuration measured on the NVIDIA forum), then after it is stable switch to &lt;code&gt;-cc.mode=0 -cc.cudagraph_mode=FULL_DECODE_ONLY&lt;/code&gt;: skip inductor compilation that hangs on GB10, keep only decode CUDA graph. Production is already the latter. Note that the field name must be &lt;strong&gt;underscore&lt;/strong&gt; &lt;code&gt;cudagraph_mode&lt;/code&gt;; writing it with a hyphen will be rejected by pydantic&lt;/p&gt;

&lt;h3&gt;
  
  
  Startup and watchdog
&lt;/h3&gt;

&lt;p&gt;Before starting, drop page cache on both machines. When BF16 failed previously, &lt;code&gt;buff/cache&lt;/code&gt; had already reached 110 GiB; the page cache from reading safetensors and the driver were fighting over the same unified memory:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sync&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;echo &lt;/span&gt;3 | &lt;span class="nb"&gt;sudo tee&lt;/span&gt; /proc/sys/vm/drop_caches
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Start rank 0 first, then rank 1:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Local&lt;/span&gt;
&lt;span class="nb"&gt;cd&lt;/span&gt; /home/lenovo/vllm
&lt;span class="nb"&gt;sudo &lt;/span&gt;docker compose &lt;span class="nt"&gt;-f&lt;/span&gt; qwen3.8-flash-next-compose.yml &lt;span class="nt"&gt;--env-file&lt;/span&gt; .env up &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--force-recreate&lt;/span&gt; &lt;span class="nt"&gt;--no-deps&lt;/span&gt;

&lt;span class="c"&gt;# Peer, via QSFP&lt;/span&gt;
ssh &lt;span class="nt"&gt;-T&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;Compression&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;no &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;BatchMode&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;yes&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; aes128-gcm@openssh.com &lt;span class="se"&gt;\&lt;/span&gt;
  lenovo@169.254.159.206 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s1"&gt;'cd /home/lenovo/vllm &amp;amp;&amp;amp; sudo docker compose -f qwen3.8-flash-next-compose.yml --env-file .env up -d --force-recreate --no-deps'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Do not use&lt;/strong&gt; &lt;code&gt;nvidia-smi&lt;/code&gt; &lt;strong&gt;for liveness probes.&lt;/strong&gt; Once the driver is stuck in &lt;code&gt;nv_alloc_system_pages&lt;/code&gt; holding the RM global write lock, &lt;code&gt;nvidia-smi&lt;/code&gt; itself will also be blocked by that same lock—in the BF16 incident it blocked for more than 614 seconds. The correct liveness probe is polling &lt;code&gt;MemAvailable&lt;/code&gt; in &lt;code&gt;/proc/meminfo&lt;/code&gt;: if it drops below &lt;strong&gt;5 GiB&lt;/strong&gt;, immediately &lt;code&gt;docker stop&lt;/code&gt;, seizing disposal rights before the kernel. The watchdog script is in the appendix.&lt;/p&gt;

&lt;p&gt;Expected log order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;rank 1: &lt;code&gt;Launching vLLM ... headless multiproc executor&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Both sides: &lt;code&gt;rank N in world size 2 is assigned as ... TP rank N&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Loading model from scratch...&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Model loading took 86.33 GiB memory ...&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;GPU KV cache size: ... tokens&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;rank 0: &lt;code&gt;Graph capturing finished in 5 secs, took 0.36 GiB&lt;/code&gt; (this line is absent in eager mode)&lt;/li&gt;
&lt;li&gt;rank 0: &lt;code&gt;Application startup complete.&lt;/code&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Timeline of this CUDA graph round:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;Peer&lt;/th&gt;
&lt;th&gt;Local&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Load weights&lt;/td&gt;
&lt;td&gt;300.79 s (model ready 307.62 s)&lt;/td&gt;
&lt;td&gt;850.66 s (model ready 857.87 s)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;init engine (profile / KV / warmup)&lt;/td&gt;
&lt;td&gt;Aligned with local, stuck on collective communication&lt;/td&gt;
&lt;td&gt;73.57 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CUDA graph capture&lt;/td&gt;
&lt;td&gt;Counted in the previous row&lt;/td&gt;
&lt;td&gt;5 s / 0.36 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MemAvailable after startup&lt;/td&gt;
&lt;td&gt;about 15 GiB&lt;/td&gt;
&lt;td&gt;about 12 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;During loading, MemAvailable dropping from 110+ GiB to 17–24 GiB is normal; the weights have entered unified memory. The watchdog threshold of 5 GiB still leaves headroom. Local loading is clearly slower than the peer; the root cause is that the local root partition was fuller at the time, and sequentially reading 173G of safetensors eats page cache more; &lt;code&gt;--distributed-timeout-seconds 3600&lt;/code&gt; was prepared for this skew&lt;/p&gt;

&lt;h3&gt;
  
  
  Verification
&lt;/h3&gt;

&lt;p&gt;Only hit rank 0. First check whether the model is registered:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; http://localhost:20001/v1/models
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It returns &lt;code&gt;id&lt;/code&gt; as &lt;code&gt;qwen3.8-flash-next&lt;/code&gt;, &lt;code&gt;root&lt;/code&gt; as &lt;code&gt;Qwen/Qwen3.8-Flash-Next-FP8&lt;/code&gt;, &lt;code&gt;max_model_len&lt;/code&gt; as 262144.&lt;/p&gt;

&lt;p&gt;Token generation and thinking. Qwen3.8 defaults to &lt;code&gt;enable_thinking=True&lt;/code&gt;, and vLLM has &lt;code&gt;--reasoning-parser qwen3&lt;/code&gt; enabled. Actual test of "factorial of 3": &lt;code&gt;completion_tokens=207&lt;/code&gt;, of which &lt;code&gt;reasoning_tokens=189&lt;/code&gt;, and the body gives "3×2×1=6". This version of the API puts the thinking process in &lt;code&gt;message.reasoning&lt;/code&gt; / &lt;code&gt;usage.completion_tokens_details.reasoning_tokens&lt;/code&gt;; the old field &lt;code&gt;reasoning_content&lt;/code&gt; may not still be present.&lt;/p&gt;

&lt;p&gt;Tool calling. &lt;code&gt;--enable-auto-tool-choice --tool-call-parser qwen3_xml&lt;/code&gt;, given a &lt;code&gt;get_weather&lt;/code&gt; function schema, asked "How is the weather now? Please call the tool to query":&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;finish_reason&lt;/code&gt;: &lt;code&gt;tool_calls&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;tool_calls[0].function.name&lt;/code&gt;: &lt;code&gt;get_weather&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;arguments&lt;/code&gt;: &lt;code&gt;{"city": "Beijing"}&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Only when all three of these pass can the service be considered usable, rather than "the process is still there and the port can be connected"&lt;/p&gt;

&lt;h3&gt;
  
  
  Postmortem of pitfalls
&lt;/h3&gt;

&lt;p&gt;This is the part the whole article really needs to leave behind; all three pitfalls exploded on the production path&lt;/p&gt;

&lt;h4&gt;
  
  
  BF16: not OOM, but whole-machine deadlock
&lt;/h4&gt;

&lt;p&gt;Following the order of "try BF16 first, then FP8 if GPU memory is insufficient", after BF16 loaded to &lt;code&gt;Loading model from scratch...&lt;/code&gt;, not a single compute kernel ever started, GPU utilization 0%, yet the CPU was saturated by kernel mode. The call stack from the kernel log (boot -1):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;VLLM::Worker_TP:507121 &amp;lt;writer&amp;gt;  holds RM global rw-semaphore
  nvidia_unlocked_ioctl -&amp;gt; RmIoctl -&amp;gt; Nv04AllocWithAccessSecInfo
  -&amp;gt; rmapiAllocWithSecInfo -&amp;gt; memdescAlloc -&amp;gt; osAllocPagesInternal
  -&amp;gt; nv_alloc_pages -&amp;gt; nv_alloc_system_pages   &amp;lt;-- stuck here
INFO: task nvidia-smi:508972 blocked for more than 614 seconds
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Under unified memory, allocating 167.6 GiB of weights is asking the kernel for 167.6 GiB of system pages, while physically there are only 119 GiB. The driver cannot get the pages, so it repeatedly reclaims and retries while holding the RM write lock. The OOM killer never managed to intervene—it kills user-mode processes, and cannot kill a kernel path stuck in the driver allocation loop. At the time &lt;code&gt;buff/cache&lt;/code&gt; was already 110 GiB; the page cache built up by reading weights and the driver fought over the same memory, aggravating the deadlock. In the end the local machine could only be rebooted.&lt;br&gt;
&lt;/p&gt;

&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart TD
  Load["vLLM loads weights"] --&amp;gt; Need{"Per-node demand vs 119 GiB"}
  Need --&amp;gt;|"BF16 167.6 GiB exceeds"| Spin["nv_alloc_system_pages lock-held spin"]
  Spin --&amp;gt; Dead["CPU kernel mode saturated / GPU 0% / whole-machine reboot"]
  Need --&amp;gt;|"FP8 86.4 GiB fits"| Ok["Allocation succeeds -&amp;gt; KV cache -&amp;gt; service ready"]
  Spin -.-&amp;gt;|"added memory: 105G"| Killed["cgroup reclaims page cache, kills only the container if over limit"]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;The point of the guardrail is here: the 105G memcg cap turns over-limit into "kill only the container". It cannot save BF16—167 GiB will never fit into 119 GiB—but it can stop the next configuration mistake from dragging the whole machine to death. &lt;code&gt;restart: "no"&lt;/code&gt; guarantees that after being killed it will not automatically charge in again&lt;/p&gt;

&lt;h4&gt;
  
  
  rank 1 omitted &lt;code&gt;--headless&lt;/code&gt;
&lt;/h4&gt;

&lt;p&gt;In the first FP8 round both sides ran a full EngineCore. The weights had actually already been loaded in: peer 340.86 s, local 831.18 s, MemAvailable dropped to 17–23 GiB, looking quite healthy. Then on the peer in &lt;code&gt;_initialize_kv_caches&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AssertionError: collective_rpc should not be called on follower node
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once the peer died, the local NCCL followed with &lt;code&gt;IBV_WC_RETRY_EXC_ERR&lt;/code&gt;. In the image source, &lt;code&gt;--headless&lt;/code&gt; is commented as “headless workers (for multi-node PP/TP)”. After compose used bash to add this argument to rank 1 according to &lt;code&gt;NODE_RANK&lt;/code&gt;, KV initialization passed on the first try&lt;/p&gt;

&lt;h4&gt;
  
  
  The CUDA graph field name uses underscores
&lt;/h4&gt;

&lt;p&gt;The plan wrote &lt;code&gt;--cudagraph-mode FULL_DECODE_ONLY&lt;/code&gt;. This version of the CLI does not have this top-level argument; the nested form &lt;code&gt;-cc.cudagraph-mode=FULL_DECODE_ONLY&lt;/code&gt; is passed as-is by FlexibleArgumentParser to pydantic, which reports:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;argument --compilation-config/-cc: 1 validation error for CompilationConfig
cudagraph-mode
  Unexpected keyword argument
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The container immediately exit 2. The correct form:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;-&lt;span class="n"&gt;cc&lt;/span&gt;.&lt;span class="n"&gt;mode&lt;/span&gt;=&lt;span class="m"&gt;0&lt;/span&gt;
-&lt;span class="n"&gt;cc&lt;/span&gt;.&lt;span class="n"&gt;cudagraph_mode&lt;/span&gt;=&lt;span class="n"&gt;FULL_DECODE_ONLY&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Equivalent JSON also works: &lt;code&gt;-cc '{"mode":0,"cudagraph_mode":"FULL_DECODE_ONLY"}'&lt;/code&gt;. What the logs confirm as taking effect is &lt;code&gt;cudagraph_mode: &amp;lt;CUDAGraphMode.FULL_DECODE_ONLY: (2, 0)&amp;gt;&lt;/code&gt;, capture sizes &lt;code&gt;[1, 2, 4, 8, 16]&lt;/code&gt; (constrained by &lt;code&gt;--max-num-seqs 8&lt;/code&gt;)&lt;/p&gt;

&lt;h3&gt;
  
  
  Performance and trade-offs
&lt;/h3&gt;

&lt;p&gt;The same request to "count from 1 to 80", &lt;code&gt;temperature=0&lt;/code&gt;, &lt;code&gt;enable_thinking=False&lt;/code&gt;, &lt;code&gt;completion_tokens=231&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Configuration&lt;/th&gt;
&lt;th&gt;Time&lt;/th&gt;
&lt;th&gt;tok/s&lt;/th&gt;
&lt;th&gt;CUDA graph occupancy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;--enforce-eager&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;10.211 s&lt;/td&gt;
&lt;td&gt;22.62&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;-cc.mode=0&lt;/code&gt; + &lt;code&gt;FULL_DECODE_ONLY&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;9.905 s&lt;/td&gt;
&lt;td&gt;23.32&lt;/td&gt;
&lt;td&gt;local 0.36 GiB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The community has reported order-of-magnitude decode speedups on this architecture; we did not see that. The reason is very specific: decode is already fast, and the cross-machine TP RoCE allreduce is the main path; CUDA graph can only eat the local kernel-launch overhead, not that trip over the network. The graph only occupies an extra 0.36 GiB, and KV went from eager's 503,285 tokens / 1.92x to 541,401 tokens / 2.07x, without blowing the budget, so production keeps it.&lt;/p&gt;

&lt;p&gt;Do not enable the graph before you have stood firm. Eager is the baseline for verifying correctness; the graph is a bonus after correctness&lt;/p&gt;

&lt;h1&gt;
  
  
  Closing
&lt;/h1&gt;

&lt;p&gt;This article spans a long stretch; from repairing the PGX to building the dual-GPU large model, about two months passed in between, but fortunately in the end all the pitfalls were stepped in, and the tool also went live (in the middle the whole machine locked up and I once thought I would have to repeat the "reinstall the OS" step); the configuration itself is not hard, what matters most is the process of troubleshooting. The IP addresses and so on mentioned earlier need to be changed to the corresponding content on your own machines; there is quite a bit of content in the article, and omissions are hard to avoid, so please correct me, readers&lt;/p&gt;

&lt;h1&gt;
  
  
  Appendix A: the two &lt;code&gt;.env&lt;/code&gt; files
&lt;/h1&gt;

&lt;p&gt;Local &lt;code&gt;/home/lenovo/vllm/.env&lt;/code&gt; (rank 0):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;MODEL_ID&lt;/span&gt;=&lt;span class="n"&gt;Qwen&lt;/span&gt;/&lt;span class="n"&gt;Qwen3&lt;/span&gt;.&lt;span class="m"&gt;8&lt;/span&gt;-&lt;span class="n"&gt;Flash&lt;/span&gt;-&lt;span class="n"&gt;Next&lt;/span&gt;-&lt;span class="n"&gt;FP8&lt;/span&gt;
&lt;span class="n"&gt;NODE_RANK&lt;/span&gt;=&lt;span class="m"&gt;0&lt;/span&gt;
&lt;span class="n"&gt;VLLM_HOST_IP&lt;/span&gt;=&lt;span class="m"&gt;169&lt;/span&gt;.&lt;span class="m"&gt;254&lt;/span&gt;.&lt;span class="m"&gt;94&lt;/span&gt;.&lt;span class="m"&gt;252&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same-named file on the peer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;&lt;span class="n"&gt;MODEL_ID&lt;/span&gt;=&lt;span class="n"&gt;Qwen&lt;/span&gt;/&lt;span class="n"&gt;Qwen3&lt;/span&gt;.&lt;span class="m"&gt;8&lt;/span&gt;-&lt;span class="n"&gt;Flash&lt;/span&gt;-&lt;span class="n"&gt;Next&lt;/span&gt;-&lt;span class="n"&gt;FP8&lt;/span&gt;
&lt;span class="n"&gt;NODE_RANK&lt;/span&gt;=&lt;span class="m"&gt;1&lt;/span&gt;
&lt;span class="n"&gt;VLLM_HOST_IP&lt;/span&gt;=&lt;span class="m"&gt;169&lt;/span&gt;.&lt;span class="m"&gt;254&lt;/span&gt;.&lt;span class="m"&gt;159&lt;/span&gt;.&lt;span class="m"&gt;206&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;MASTER_ADDR&lt;/code&gt; is written in the compose, fixed to the local &lt;code&gt;169.254.94.252&lt;/code&gt;, and should not change along with &lt;code&gt;VLLM_HOST_IP&lt;/code&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  Appendix B: the complete compose
&lt;/h1&gt;

&lt;p&gt;File path: &lt;code&gt;/home/lenovo/vllm/qwen3.8-flash-next-compose.yml&lt;/code&gt;. The content on both machines must be identical; after changing it, push to the peer via QSFP with scp.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Dual-node TP=2 for Qwen3.8-Flash-Next across two GB10 workstations.&lt;/span&gt;
&lt;span class="c1"&gt;# Same file on both machines; NODE_RANK / VLLM_HOST_IP / MODEL_ID come from .env.&lt;/span&gt;
&lt;span class="c1"&gt;# host network: vLLM binds :20001 directly (same external port as qwen3.6-35b-compose.yml).&lt;/span&gt;
&lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;vllm&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;vllm/vllm-openai:qwen38-flash-next&lt;/span&gt;
    &lt;span class="na"&gt;container_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;vllm-qwen38&lt;/span&gt;
    &lt;span class="na"&gt;network_mode&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;host&lt;/span&gt;
    &lt;span class="na"&gt;ipc&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;host&lt;/span&gt;
    &lt;span class="na"&gt;restart&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;no"&lt;/span&gt;
    &lt;span class="na"&gt;stdin_open&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
    &lt;span class="na"&gt;tty&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
    &lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;/home/lenovo/vllm/huggingface:/root/.cache/huggingface&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;/etc/localtime:/etc/localtime:ro&lt;/span&gt;
    &lt;span class="na"&gt;devices&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;/dev/infiniband:/dev/infiniband&lt;/span&gt;
    &lt;span class="na"&gt;cap_add&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;IPC_LOCK&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;SYS_PTRACE&lt;/span&gt;
    &lt;span class="na"&gt;security_opt&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;seccomp:unconfined&lt;/span&gt;
    &lt;span class="na"&gt;ulimits&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;memlock&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;-1&lt;/span&gt;
      &lt;span class="na"&gt;stack&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;67108864&lt;/span&gt;
    &lt;span class="na"&gt;deploy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;limits&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;105G&lt;/span&gt;
        &lt;span class="na"&gt;reservations&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;devices&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
            &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;driver&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;nvidia&lt;/span&gt;
              &lt;span class="na"&gt;count&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;all&lt;/span&gt;
              &lt;span class="na"&gt;capabilities&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;gpu&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;PYTHONUNBUFFERED&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1"&lt;/span&gt;
      &lt;span class="na"&gt;MODEL_ID&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${MODEL_ID}&lt;/span&gt;
      &lt;span class="na"&gt;NODE_RANK&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${NODE_RANK}&lt;/span&gt;
      &lt;span class="na"&gt;VLLM_HOST_IP&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${VLLM_HOST_IP}&lt;/span&gt;
      &lt;span class="na"&gt;MASTER_ADDR&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;169.254.94.252"&lt;/span&gt;
      &lt;span class="na"&gt;NCCL_IB_HCA&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;=rocep1s0f0,roceP2p1s0f0"&lt;/span&gt;
      &lt;span class="na"&gt;NCCL_IB_DISABLE&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0"&lt;/span&gt;
      &lt;span class="na"&gt;NCCL_IB_GID_INDEX&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;3"&lt;/span&gt;
      &lt;span class="na"&gt;NCCL_IB_MERGE_NICS&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1"&lt;/span&gt;
      &lt;span class="na"&gt;NCCL_SOCKET_IFNAME&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;enp1s0f0np0&lt;/span&gt;
      &lt;span class="na"&gt;NCCL_SOCKET_FAMILY&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;AF_INET&lt;/span&gt;
      &lt;span class="na"&gt;GLOO_SOCKET_IFNAME&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;enp1s0f0np0&lt;/span&gt;
      &lt;span class="na"&gt;GLOO_SOCKET_FAMILY&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;AF_INET&lt;/span&gt;
      &lt;span class="na"&gt;GLOO_USE_LIBUV&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0"&lt;/span&gt;
      &lt;span class="na"&gt;TORCH_GLOO_USE_LIBUV&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0"&lt;/span&gt;
      &lt;span class="na"&gt;TP_SOCKET_IFNAME&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;enp1s0f0np0&lt;/span&gt;
      &lt;span class="na"&gt;NCCL_NET_PLUGIN&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;none&lt;/span&gt;
      &lt;span class="na"&gt;NCCL_IB_ROCE_VERSION_NUM&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2"&lt;/span&gt;
      &lt;span class="na"&gt;NCCL_CUMEM_ENABLE&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;0"&lt;/span&gt;
      &lt;span class="na"&gt;NCCL_DEBUG&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;WARN&lt;/span&gt;
    &lt;span class="c1"&gt;# Rank 1 must be --headless (worker only). Running a full EngineCore on the&lt;/span&gt;
    &lt;span class="c1"&gt;# follower hits: AssertionError: collective_rpc should not be called on follower node&lt;/span&gt;
    &lt;span class="na"&gt;entrypoint&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/bin/bash"&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;-lc"&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="na"&gt;command&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
        &lt;span class="s"&gt;set -euo pipefail&lt;/span&gt;
        &lt;span class="s"&gt;if [ "${NODE_RANK}" = "1" ]; then&lt;/span&gt;
          &lt;span class="s"&gt;exec vllm serve "${MODEL_ID}" \&lt;/span&gt;
            &lt;span class="s"&gt;--served-model-name qwen3.8-flash-next \&lt;/span&gt;
            &lt;span class="s"&gt;--tensor-parallel-size 2 \&lt;/span&gt;
            &lt;span class="s"&gt;--nnodes 2 \&lt;/span&gt;
            &lt;span class="s"&gt;--node-rank "${NODE_RANK}" \&lt;/span&gt;
            &lt;span class="s"&gt;--master-addr 169.254.94.252 \&lt;/span&gt;
            &lt;span class="s"&gt;--master-port 29501 \&lt;/span&gt;
            &lt;span class="s"&gt;--distributed-executor-backend mp \&lt;/span&gt;
            &lt;span class="s"&gt;-cc.mode=0 \&lt;/span&gt;
            &lt;span class="s"&gt;-cc.cudagraph_mode=FULL_DECODE_ONLY \&lt;/span&gt;
            &lt;span class="s"&gt;--gpu-memory-utilization 0.80 \&lt;/span&gt;
            &lt;span class="s"&gt;--distributed-timeout-seconds 3600 \&lt;/span&gt;
            &lt;span class="s"&gt;--max-model-len 262144 \&lt;/span&gt;
            &lt;span class="s"&gt;--max-num-seqs 8 \&lt;/span&gt;
            &lt;span class="s"&gt;--enable-prefix-caching \&lt;/span&gt;
            &lt;span class="s"&gt;--no-enable-flashinfer-autotune \&lt;/span&gt;
            &lt;span class="s"&gt;--reasoning-parser qwen3 \&lt;/span&gt;
            &lt;span class="s"&gt;--enable-auto-tool-choice \&lt;/span&gt;
            &lt;span class="s"&gt;--tool-call-parser qwen3_xml \&lt;/span&gt;
            &lt;span class="s"&gt;--host 0.0.0.0 \&lt;/span&gt;
            &lt;span class="s"&gt;--port 20001 \&lt;/span&gt;
            &lt;span class="s"&gt;--headless&lt;/span&gt;
        &lt;span class="s"&gt;fi&lt;/span&gt;
        &lt;span class="s"&gt;exec vllm serve "${MODEL_ID}" \&lt;/span&gt;
          &lt;span class="s"&gt;--served-model-name qwen3.8-flash-next \&lt;/span&gt;
          &lt;span class="s"&gt;--tensor-parallel-size 2 \&lt;/span&gt;
          &lt;span class="s"&gt;--nnodes 2 \&lt;/span&gt;
          &lt;span class="s"&gt;--node-rank "${NODE_RANK}" \&lt;/span&gt;
          &lt;span class="s"&gt;--master-addr 169.254.94.252 \&lt;/span&gt;
          &lt;span class="s"&gt;--master-port 29501 \&lt;/span&gt;
          &lt;span class="s"&gt;--distributed-executor-backend mp \&lt;/span&gt;
          &lt;span class="s"&gt;-cc.mode=0 \&lt;/span&gt;
          &lt;span class="s"&gt;-cc.cudagraph_mode=FULL_DECODE_ONLY \&lt;/span&gt;
          &lt;span class="s"&gt;--gpu-memory-utilization 0.80 \&lt;/span&gt;
          &lt;span class="s"&gt;--distributed-timeout-seconds 3600 \&lt;/span&gt;
          &lt;span class="s"&gt;--max-model-len 262144 \&lt;/span&gt;
          &lt;span class="s"&gt;--max-num-seqs 8 \&lt;/span&gt;
          &lt;span class="s"&gt;--enable-prefix-caching \&lt;/span&gt;
          &lt;span class="s"&gt;--no-enable-flashinfer-autotune \&lt;/span&gt;
          &lt;span class="s"&gt;--reasoning-parser qwen3 \&lt;/span&gt;
          &lt;span class="s"&gt;--enable-auto-tool-choice \&lt;/span&gt;
          &lt;span class="s"&gt;--tool-call-parser qwen3_xml \&lt;/span&gt;
          &lt;span class="s"&gt;--host 0.0.0.0 \&lt;/span&gt;
          &lt;span class="s"&gt;--port 20001&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h1&gt;
  
  
  Appendix C: watchdog script
&lt;/h1&gt;

&lt;p&gt;Run it on the local machine during loading. Do not open &lt;code&gt;nvidia-smi&lt;/code&gt; in parallel.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/usr/bin/env bash&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-euo&lt;/span&gt; pipefail

&lt;span class="nv"&gt;THRESHOLD_KB&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="k"&gt;$((&lt;/span&gt;&lt;span class="m"&gt;5&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="m"&gt;1024&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="m"&gt;1024&lt;/span&gt;&lt;span class="k"&gt;))&lt;/span&gt;
&lt;span class="nv"&gt;SSH&lt;/span&gt;&lt;span class="o"&gt;=(&lt;/span&gt;ssh &lt;span class="nt"&gt;-T&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;Compression&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;no &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;BatchMode&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;yes&lt;/span&gt; &lt;span class="nt"&gt;-c&lt;/span&gt; aes128-gcm@openssh.com lenovo@169.254.159.206&lt;span class="o"&gt;)&lt;/span&gt;

emergency_stop&lt;span class="o"&gt;()&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt;
  &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"WATCHDOG STOP: &lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
  &lt;span class="nb"&gt;sudo &lt;/span&gt;docker stop &lt;span class="nt"&gt;-t&lt;/span&gt; 10 vllm-qwen38 2&amp;gt;/dev/null &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;true&lt;/span&gt;
  &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;SSH&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s1"&gt;'sudo docker stop -t 10 vllm-qwen38 2&amp;gt;/dev/null || true'&lt;/span&gt;
&lt;span class="o"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;for &lt;/span&gt;i &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;seq &lt;/span&gt;1 210&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;&lt;span class="nv"&gt;local_avail&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="s1"&gt;'/MemAvailable/{print $2}'&lt;/span&gt; /proc/meminfo&lt;span class="si"&gt;)&lt;/span&gt;
  &lt;span class="nv"&gt;peer_avail&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;SSH&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"awk '/MemAvailable/{print &lt;/span&gt;&lt;span class="se"&gt;\$&lt;/span&gt;&lt;span class="s2"&gt;2}' /proc/meminfo"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
  &lt;span class="nv"&gt;local_status&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;docker inspect vllm-qwen38 &lt;span class="nt"&gt;--format&lt;/span&gt; &lt;span class="s1"&gt;'{{.State.Status}} {{.State.OOMKilled}} {{.State.ExitCode}}'&lt;/span&gt; 2&amp;gt;/dev/null &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;echo &lt;/span&gt;missing&lt;span class="si"&gt;)&lt;/span&gt;
  &lt;span class="nv"&gt;peer_status&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;SSH&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"sudo docker inspect vllm-qwen38 --format '{{.State.Status}} {{.State.OOMKilled}} {{.State.ExitCode}}' 2&amp;gt;/dev/null || echo missing"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
  &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"[&lt;/span&gt;&lt;span class="nv"&gt;$i&lt;/span&gt;&lt;span class="s2"&gt;] local &lt;/span&gt;&lt;span class="k"&gt;$((&lt;/span&gt;local_avail/1024/1024&lt;span class="k"&gt;))&lt;/span&gt;&lt;span class="s2"&gt;GiB &lt;/span&gt;&lt;span class="nv"&gt;$local_status&lt;/span&gt;&lt;span class="s2"&gt; | peer &lt;/span&gt;&lt;span class="k"&gt;$((&lt;/span&gt;peer_avail/1024/1024&lt;span class="k"&gt;))&lt;/span&gt;&lt;span class="s2"&gt;GiB &lt;/span&gt;&lt;span class="nv"&gt;$peer_status&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

  &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$local_avail&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-lt&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$THRESHOLD_KB&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$peer_avail&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-lt&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$THRESHOLD_KB&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;emergency_stop &lt;span class="s2"&gt;"MemAvailable below 5 GiB"&lt;/span&gt;
    &lt;span class="nb"&gt;exit &lt;/span&gt;2
  &lt;span class="k"&gt;fi
  if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$local_status&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; running&lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"LOCAL CONTAINER NOT RUNNING: &lt;/span&gt;&lt;span class="nv"&gt;$local_status&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
    &lt;span class="nb"&gt;sudo &lt;/span&gt;docker logs vllm-qwen38 2&amp;gt;&amp;amp;1 | &lt;span class="nb"&gt;tail&lt;/span&gt; &lt;span class="nt"&gt;-60&lt;/span&gt;
    &lt;span class="nb"&gt;exit &lt;/span&gt;3
  &lt;span class="k"&gt;fi
  if&lt;/span&gt; &lt;span class="o"&gt;[[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$peer_status&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; running&lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="o"&gt;]]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"PEER CONTAINER NOT RUNNING: &lt;/span&gt;&lt;span class="nv"&gt;$peer_status&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
    &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;SSH&lt;/span&gt;&lt;span class="p"&gt;[@]&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"sudo docker logs vllm-qwen38 2&amp;gt;&amp;amp;1 | tail -60"&lt;/span&gt;
    &lt;span class="nb"&gt;exit &lt;/span&gt;3
  &lt;span class="k"&gt;fi
  if &lt;/span&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;docker logs vllm-qwen38 2&amp;gt;&amp;amp;1 | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-q&lt;/span&gt; &lt;span class="s1"&gt;'Application startup complete'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'STARTUP COMPLETE'&lt;/span&gt;
    &lt;span class="nb"&gt;exit &lt;/span&gt;0
  &lt;span class="k"&gt;fi
  &lt;/span&gt;&lt;span class="nb"&gt;sleep &lt;/span&gt;10
&lt;span class="k"&gt;done

&lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'WATCHDOG TIMEOUT'&lt;/span&gt;
&lt;span class="nb"&gt;exit &lt;/span&gt;5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;transformers will print two &lt;code&gt;[ERROR] min_frames / max_frames ... not documented&lt;/code&gt; lines; that is docstring noise, do not treat it as a startup failure&lt;/p&gt;

&lt;h1&gt;
  
  
  Appendix D: troubleshooting
&lt;/h1&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symptom&lt;/th&gt;
&lt;th&gt;Action&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Logs stuck on IPv6 / Gloo timeout&lt;/td&gt;
&lt;td&gt;Confirm&lt;code&gt;GLOO_USE_LIBUV=0&lt;/code&gt;, &lt;code&gt;NCCL_SOCKET_FAMILY=AF_INET&lt;/code&gt;, and UFW has allowed 169.254/16 and 192.168.101.0/24&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;collective_rpc should not be called on follower node&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;rank 1 must have&lt;code&gt;--headless&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;cudagraph-mode Unexpected keyword argument&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Use&lt;code&gt;-cc.cudagraph_mode&lt;/code&gt;, not hyphens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;CUDA out of memory&lt;/code&gt; / insufficient KV cache&lt;/td&gt;
&lt;td&gt;Lower&lt;code&gt;--max-model-len&lt;/code&gt; along 262144 → 131072 → 65536&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mamba cache capacity error&lt;/td&gt;
&lt;td&gt;Raise&lt;code&gt;--max-num-seqs&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CPU saturated, GPU 0%,&lt;code&gt;nvidia-smi&lt;/code&gt; stuck&lt;/td&gt;
&lt;td&gt;Immediately&lt;code&gt;docker stop&lt;/code&gt;, do not wait further; check whether local weights were switched back to BF16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MemAvailable drops below 5 GiB&lt;/td&gt;
&lt;td&gt;The watchdog should already have stopped the containers; review before starting again, do not use&lt;code&gt;restart: always&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPU not visible inside Docker on the peer&lt;/td&gt;
&lt;td&gt;Check driver kernel-mode / user-mode versions; if they do not match, reboot the peer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;rsync fails on&lt;code&gt;trees/*.json&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;sudo chmod -R a+rX&lt;/code&gt; the corresponding hub directory before transferring again&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Example chat request for verification:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; http://localhost:20001/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "qwen3.8-flash-next",
    "messages": [{"role": "user", "content": "Answer in one sentence: what is the factorial of 3?"}],
    "max_tokens": 256,
    "chat_template_kwargs": {"enable_thinking": true}
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



</description>
      <category>ai</category>
    </item>
  </channel>
</rss>
