DEV Community

Ankit Khandelwal
Ankit Khandelwal

Posted on

Running LLMs on AMD NPU with OpenFlowLM - Fedora Guide

Tested on Fedora 44, kernel 7.2.7-200.fc44.x86_64,
ROG Flow Z13 (Ryzen AI Max 390 / Strix Halo NPU).

Goal: copy-paste setup that gets oflm validate and oflm run working from
a source build, with a Python setup that touches nothing system-wide.

Companion piece: the FastFlowLM Fedora guide
covers the driver + XRT layers. This guide assumes those already work
(flm validate passes) and builds OpenFlowLM on top.

TL;DR

You need the same four layers as FastFlowLM, with OpenFlowLM replacing
flm:

Layer What it does
Kernel + DKMS driver (amdxdna) Creates /dev/accel/accel0, loads NPU firmware
XRT base AMD runtime in /opt/xilinx/xrt
XRT NPU plugin (xrt_plugin RPM) Provides libxrt_driver_xdna.so so XRT sees the NPU
OpenFlowLM (oflm) Open-kernel LLM runtime, built from source here

Two OpenFlowLM-specific gotchas beyond the FLM guide:

  1. The kernel-toolchain Python must match your XRT's pyxrt. utilities/export-kernels.py pins the ironvenv Python via PYTHON =. A source-built XRT ships pyxrt for one Python only — check /opt/xilinx/xrt/python/ for the cpython-*.so tag and set the pin to match (on this host: 3.14).
  2. Python tool deps live in a venv, never in system Python. q4nx-build needs torch + friends; install CPU-only torch from the PyTorch CPU index so you don't pull multi-GB CUDA blobs onto an NPU box.

Hardware tested

Item Value
Machine ASUS ROG Flow Z13 GZ302EA
CPU / NPU AMD Ryzen AI Max 390 (Strix Halo)
NPU PCI ID 1022:17f0 rev 11
OS Fedora Linux 44 Workstation
Kernel 7.2.7-200.fc44.x86_64
NPU firmware 1.1.2.65
XRT 2.25.0 (built from amd/xdna-driver)
Toolchain CMake 4.3, Ninja 1.13, GCC 16, Python 3.14

0. Confirm the base stack (from the FLM guide)

cat /proc/cmdline | grep amd_iommu || echo "OK: amd_iommu not disabled"
ls -la /dev/accel/
ulimit -l   # want: unlimited
/opt/xilinx/xrt/bin/xrt-smi examine   # want: NPU Strix Halo, 6x8
flm validate
Enter fullscreen mode Exit fullscreen mode

If any of that fails, work through the FLM guide first — oflm uses the
same kernel device and the same XRT.

1. Clone (recursive — don't skip this)

git clone --recursive https://github.com/Atomic-Germ/OpenFlowLM-Next.git ~/repos/OpenFlowLM-Next
cd ~/repos/OpenFlowLM-Next
Enter fullscreen mode Exit fullscreen mode

The third_party/tokenizers-cpp submodule (plus its nested sentencepiece
/ msgpack) is required at configure time. If you cloned without
--recursive:

git submodule update --init --recursive
Enter fullscreen mode Exit fullscreen mode

2. Match the ironvenv Python pin to your XRT

The full build auto-creates ironvenv/ (mlir-aie 1.4.2 + Peano llvm-aie) and
uses it to compile every open-kernel xclbin. The BERT embedding export also
imports pyxrt from /opt/xilinx/xrt/python — which is built for exactly
one Python version. Check yours:

ls /opt/xilinx/xrt/python/
# e.g. pyxrt.cpython-314-x86_64-linux-gnu.so  -> pin is "3.14"
Enter fullscreen mode Exit fullscreen mode

Then make utilities/export-kernels.py match:

# pyxrt is built for 3.14 by this host's XRT (see /opt/xilinx/xrt/python);
# pin the venv to match. Adjust if your XRT ships a different cpython-*.so.
PYTHON = "3.14"
Enter fullscreen mode Exit fullscreen mode

Mismatch symptom: import pyxrt / No module named 'pyxrt' (or a
cpython-311 vs cpython-314 ABI error) halfway through the kernel export,
after the compile-only specs already succeeded.

3. Tools venv (q4nx-build + oflm-test, no system changes)

Everything Python outside the kernel toolchain lives in one user-owned venv.
No sudo pip, no system site-packages:

/usr/bin/python3.14 -m venv ~/.venvs/oflm-tools
V=~/.venvs/oflm-tools
$V/bin/pip install --upgrade pip
# CPU-only torch first: the default index pulls GBs of NVIDIA CUDA wheels
# that an NPU-only box will never use. Install torch from the CPU index,
# then the editable packages resolve everything else around it.
$V/bin/pip install --index-url https://download.pytorch.org/whl/cpu torch
$V/bin/pip install -e ./utilities/q4nx-build
$V/bin/pip install -e ./utilities/oflm-test
Enter fullscreen mode Exit fullscreen mode

Editable (-e) installs mean q4nx / oflm-test source edits take effect
immediately — the venv for feature work, not just running. Verify:

$V/bin/q4nx-build --help
$V/bin/oflm-test --help
Enter fullscreen mode Exit fullscreen mode

4. Build: engine-only first, full second

Engine-only (fast iteration, ~minutes on 24 cores). No ironvenv, no
NPU needed at build time; runs models using the in-tree src/xclbins/ sets:

cmake --preset linux-debug
cmake --build --preset linux-debug
ctest --test-dir build-debug --output-on-failure
Enter fullscreen mode Exit fullscreen mode

Full distribution (engine + all open-kernel xclbins). Kicks off the
export_kernels target: creates ironvenv, compiles every spec in
open_kernels/recipes/specs/, then builds the BERT design sets on the NPU.
Budget well over an hour; run it detached because ninja buffers the export
script's output until the target finishes:

cmake --preset linux-default
setsid nohup cmake --build --preset linux-default > /tmp/oflm-full-build.log 2>&1 < /dev/null &
# poll with: ps aux | grep "[e]xport_qwen" (bracket avoids matching your own grep;
# plain pgrep -f matches its own command line and lies to you)
ctest --preset linux-default   # smoke: oflm list + openai_compat + vision + embed
Enter fullscreen mode Exit fullscreen mode

ninja swallows per-spec progress — while it runs, watch the real work with
ps aux | grep "[e]xport_qwen" (Peano clang++ --target=aie2p-... lines mean
AIE compiles are flowing).

Rebuilds are cheap. Re-running the full build re-passes every spec, but
each design hits the per-design build cache under open_kernels/designs/ and
~/.npu/cache, so an already-built spec verifies in seconds instead of
recompiling — a no-change rebuild takes minutes, not another hour. Only
changed/new shapes pay full compile cost, followed by whatever export stage
hadn't finished (e.g. the BERT sets). While iterating on one family, skip
the rest entirely:

cmake -B build --preset linux-default -DOFLM_KERNEL_SPECS=qwen3-4b,gemma3-4b
Enter fullscreen mode Exit fullscreen mode

5. Install

With a terminal (password prompt works):

sudo cmake --install build        # full build -> /opt/openflowlm + /usr/bin/oflm
# or: sudo cmake --install build-debug   # engine-only dev install
Enter fullscreen mode Exit fullscreen mode

Without sudo (user-space staging, same layout, no root):

DESTDIR=~/oflm-install cmake --install build
~/oflm-install/opt/openflowlm/bin/oflm validate
Enter fullscreen mode Exit fullscreen mode

Note: the staged usr/bin/oflm symlink points at absolute /opt/... and
only resolves after a real install — use opt/openflowlm/bin/oflm directly
from a staging dir.

6. Validate and run

oflm validate
# [Linux]  NPU: /dev/accel/accel0 with 8 columns
# [Linux]  NPU FW Version: 1.1.2.65
# [Linux]  Memlock Limit: infinity

oflm list
oflm run gemma4-it:e4b
oflm serve gemma4-it:e4b     # OpenAI-compatible server on port 52625
oflm bench gemma4-it:e4b
Enter fullscreen mode Exit fullscreen mode

Inside oflm run: /verbose (per-turn TTFT, prefill tok/s, decoding tok/s)
and /status (token counts, throughput summary). Models download from
HuggingFace on first run into ~/.config/oflm/ (~/.config/flm/ is still
searched first for pre-rename installs).

Full test suite from the venv:

~/.venvs/oflm-tools/bin/oflm-test --llm
~/.venvs/oflm-tools/bin/oflm-test --embedding
Enter fullscreen mode Exit fullscreen mode

7. Monitor NPU stats

Same story as FLM — no amdgpu_top equivalent for the NPU. Combine XRT
snapshots with in-session metrics:

xrt-smi examine
xrt-smi examine -r all -d 0000:c5:00.1
watch -n1 'xrt-smi examine -r all -d 0000:c5:00.1'   # poll beside a run
xrt-smi validate    # hardware sanity: gemm TOPS, latency, throughput
Enter fullscreen mode Exit fullscreen mode

Troubleshooting

Symptom Cause Fix
No module named 'pyxrt' mid-export ironvenv Python ≠ XRT's pyxrt build Section 2: match PYTHON to /opt/xilinx/xrt/python/*.so tag, rm -rf ironvenv, rebuild
FileNotFoundError: ~/.npu/cache in export_gemm_rtp.py purge Fresh machine, cache dir never created (fixed in-tree: purge/find_cache now ensure it) mkdir -p ~/.npu/cache on older checkouts, or pull the fix and rebuild
model.q4nx is N bytes, manifest says M — treating as missing Truncated download in ~/.config/flm/models/<M>/ Delete the undersized model.q4nx / 0-byte tokenizer.json; next run re-fetches
No module named 'einops' from q4nx-build Tool venv deps missing Section 3; activate/use ~/.venvs/oflm-tools
oflm-test: No module named 'oflm_test' from a DESTDIR staging dir Launcher hardcodes /opt/openflowlm paths Expected in staging — use PYTHONPATH=<staging>/.../utilities/oflm-test or do the real sudo install
No NPU device found / PASID unavailable amd_iommu=off on cmdline Remove it from GRUB, regenerate, reboot (see FLM guide §5)
libxrt_coreutil.so.2: cannot open XRT not in loader cache echo /opt/xilinx/xrt/lib64 > /etc/ld.so.conf.d/xrt.conf && sudo ldconfig
xrt-smi: unwrapped/xrt-smi: No such file Symlinked instead of wrapped Wrapper script, not symlink (FLM guide §7)
Submodule errors at configure (tokenizers-cpp missing) Non-recursive clone git submodule update --init --recursive
Build log looks stuck during kernels Ninja buffers export_kernels output Normal — poll with `ps aux \

Architecture recap

{% raw %}

┌─────────────────────────────────────────┐
│  oflm run / oflm serve                  │  ← this repo (engine + open kernels)
├─────────────────────────────────────────┤
│  libxrt_driver_xdna.so (XRT plugin)     │  ← xrt_plugin RPM
├─────────────────────────────────────────┤
│  libxrt_core.so (XRT base)              │  ← xrt-base RPM
├─────────────────────────────────────────┤
│  amdxdna.ko (DKMS kernel driver)        │  ← xrt_plugin RPM (postinst)
├─────────────────────────────────────────┤
│  NPU firmware (amdnpu/17f0_11/)       │  ← linux-firmware + plugin
├─────────────────────────────────────────┤
│  /dev/accel/accel0                      │  ← kernel DRM device node
└─────────────────────────────────────────┘
         ▲  requires IOMMU (PASID/SVA), memlock = unlimited
         │  kernels: src/xclbins (closed, in-tree) + open sets from ironvenv
Enter fullscreen mode Exit fullscreen mode

Top comments (0)