<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: aivrar</title>
    <description>The latest articles on DEV Community by aivrar (@aivrar).</description>
    <link>https://dev.to/aivrar</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4172649%2Fa4719f23-37e2-4156-9476-00c29c3305b7.png</url>
      <title>DEV Community: aivrar</title>
      <link>https://dev.to/aivrar</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aivrar"/>
    <language>en</language>
    <item>
      <title>Running vLLM natively on Windows</title>
      <dc:creator>aivrar</dc:creator>
      <pubDate>Fri, 09 Oct 2026 06:48:29 +0000</pubDate>
      <link>https://dev.to/aivrar/running-vllm-natively-on-windows-4ahm</link>
      <guid>https://dev.to/aivrar/running-vllm-natively-on-windows-4ahm</guid>
      <description>&lt;p&gt;vLLM officially supports Linux only. On Windows the usual answer is WSL or Docker. I wanted it running natively, with real CUDA kernels and an OpenAI-compatible server, and no Linux layer. So I maintain &lt;strong&gt;vllm-windows-build&lt;/strong&gt;: patches, prebuilt wheels, and portable installer scripts for running vLLM directly on Windows 10/11.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Disclosure:&lt;/strong&gt; I built this repo, and I built it with heavy AI-assisted coding (the README credits Claude). The engine is upstream &lt;a href="https://github.com/vllm-project/vllm" rel="noopener noreferrer"&gt;vLLM&lt;/a&gt;. My part is the Windows port, the packaging, and the testing. Everything below comes from the repo's README, docs, and release notes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's actually published
&lt;/h2&gt;

&lt;p&gt;There are three current options. The repo's &lt;a href="https://github.com/aivrar/vllm-windows-build/blob/master/docs/releases.md" rel="noopener noreferrer"&gt;release chooser&lt;/a&gt; puts it like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Channel&lt;/th&gt;
&lt;th&gt;vLLM&lt;/th&gt;
&lt;th&gt;Python&lt;/th&gt;
&lt;th&gt;Torch / CUDA&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;strong&gt;Stable / Latest&lt;/strong&gt; (&lt;code&gt;v0.27.1-win-cu130&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;0.27.1&lt;/td&gt;
&lt;td&gt;3.13.14&lt;/td&gt;
&lt;td&gt;2.13.0+cu130 / 13.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Regular release, not Latest (&lt;code&gt;v0.29.0-win-cu132-py314&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;0.29.0&lt;/td&gt;
&lt;td&gt;3.14.2&lt;/td&gt;
&lt;td&gt;2.13.0+cu132 / 13.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prerelease (&lt;code&gt;v0.29.0-win-cu130-rc1&lt;/code&gt;)&lt;/td&gt;
&lt;td&gt;0.29.0&lt;/td&gt;
&lt;td&gt;3.13.14&lt;/td&gt;
&lt;td&gt;2.13.0+cu130 / 13.0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All three use Triton for Windows 3.7.1. The 0.27.1 wheel has kernels for SM 7.5, 8.6, 8.9, and 12.0, so RTX 20/30/40/50 series. It also includes FlashAttention 2, the Rust frontend and tool parser, and opt-in CPU/filesystem prompt-KV offload.&lt;/p&gt;

&lt;p&gt;Why is 0.27.1 still the default? The Python 3.14 build uses the &lt;strong&gt;CPU TorchAudio wheel&lt;/strong&gt;, because no matching cu132 wheel was available. GPU audio processing and audio-model serving weren't validated there. Model inference itself still runs on CUDA.&lt;/p&gt;

&lt;p&gt;If you just want something that works, use 0.27.1.&lt;/p&gt;

&lt;h2&gt;
  
  
  Requirements
&lt;/h2&gt;

&lt;p&gt;From the README:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Windows 10 21H2 x64 minimum (22H2 or 11 recommended)&lt;/li&gt;
&lt;li&gt;NVIDIA GPU, SM 7.5 or newer, 12 GB VRAM minimum (24 GB recommended)&lt;/li&gt;
&lt;li&gt;16 GB RAM minimum (32 GB+ recommended)&lt;/li&gt;
&lt;li&gt;NVIDIA driver R580+ for CUDA 13.0&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Single GPU per process.&lt;/strong&gt; NCCL doesn't ship with PyTorch on Windows, so the patch uses a single-rank &lt;code&gt;FakeProcessGroup&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You don't need Visual Studio or the CUDA Toolkit unless you're building from source.&lt;/p&gt;

&lt;h2&gt;
  
  
  Install option 1: the portable installer (what I recommend)
&lt;/h2&gt;

&lt;p&gt;This route downloads its own Python, so nothing needs to be preinstalled.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open the release you picked and download &lt;strong&gt;Source code (zip)&lt;/strong&gt;. For 0.27.1 that's &lt;a href="https://github.com/aivrar/vllm-windows-build/releases/tag/v0.27.1-win-cu130" rel="noopener noreferrer"&gt;v0.27.1-win-cu130&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Extract it into a &lt;strong&gt;new, separate directory&lt;/strong&gt;. Don't mix 0.27.1, Python 3.13, and Python 3.14 installs in one folder.&lt;/li&gt;
&lt;li&gt;Run:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight batchfile"&gt;&lt;code&gt;&lt;span class="kd"&gt;install&lt;/span&gt;.bat
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It downloads Python, PyTorch, and the matching prebuilt vLLM and Multi-TurboQuant wheels. Downloads are pinned by exact size and SHA-256. Nothing gets compiled. If you already have a verified 0.27.1 wheel, put it in &lt;code&gt;dist-v0.27.1\&lt;/code&gt; next to &lt;code&gt;install.bat&lt;/code&gt; and the script will use it. Running &lt;code&gt;install.bat&lt;/code&gt; again repairs an incomplete install.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Start the server:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight batchfile"&gt;&lt;code&gt;&lt;span class="kd"&gt;launch&lt;/span&gt;.bat
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With no arguments it shows a model picker that scans &lt;code&gt;models\&lt;/code&gt; next to the script. Or pass a model directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight batchfile"&gt;&lt;code&gt;&lt;span class="kd"&gt;launch&lt;/span&gt;.bat &lt;span class="na"&gt;--model &lt;/span&gt;&lt;span class="kd"&gt;E&lt;/span&gt;:\models\Qwen3&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="m"&gt;14&lt;/span&gt;&lt;span class="kd"&gt;B&lt;/span&gt;&lt;span class="na"&gt;-AWQ&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="m"&gt;4&lt;/span&gt;&lt;span class="kd"&gt;bit&lt;/span&gt; &lt;span class="na"&gt;--port &lt;/span&gt;&lt;span class="m"&gt;8000&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;launch.bat&lt;/code&gt; checks the install first and reruns &lt;code&gt;install.bat&lt;/code&gt; if something is missing. It also sets &lt;code&gt;CUDA_DEVICE_ORDER=PCI_BUS_ID&lt;/code&gt;, so &lt;code&gt;--gpu-id&lt;/code&gt; matches &lt;code&gt;nvidia-smi&lt;/code&gt; ordering.&lt;/p&gt;

&lt;p&gt;If you prefer git, a default clone of &lt;code&gt;master&lt;/code&gt; gives you the 0.27.1 installer. For the Python 3.14 build:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight batchfile"&gt;&lt;code&gt;&lt;span class="kd"&gt;git&lt;/span&gt; &lt;span class="kd"&gt;clone&lt;/span&gt; &lt;span class="na"&gt;--branch &lt;/span&gt;&lt;span class="kd"&gt;v0&lt;/span&gt;.29.0&lt;span class="na"&gt;-win-cu&lt;/span&gt;&lt;span class="m"&gt;132&lt;/span&gt;&lt;span class="na"&gt;-py&lt;/span&gt;&lt;span class="m"&gt;314&lt;/span&gt; &lt;span class="kd"&gt;https&lt;/span&gt;://github.com/aivrar/vllm&lt;span class="na"&gt;-windows-build&lt;/span&gt;.git &lt;span class="kd"&gt;vllm&lt;/span&gt;&lt;span class="na"&gt;-py&lt;/span&gt;&lt;span class="m"&gt;314&lt;/span&gt;&lt;span class="na"&gt;-cu&lt;/span&gt;&lt;span class="m"&gt;132&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Install option 2: the wheel in your own venv
&lt;/h2&gt;

&lt;p&gt;For 0.27.1, the README's manual steps are:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight batchfile"&gt;&lt;code&gt;&lt;span class="kd"&gt;py&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="m"&gt;3&lt;/span&gt;.13 &lt;span class="na"&gt;-m &lt;/span&gt;&lt;span class="kd"&gt;venv&lt;/span&gt; &lt;span class="kd"&gt;venv&lt;/span&gt;
&lt;span class="kd"&gt;venv&lt;/span&gt;\Scripts\activate

&lt;span class="kd"&gt;pip&lt;/span&gt; &lt;span class="kd"&gt;install&lt;/span&gt; &lt;span class="kd"&gt;torch&lt;/span&gt;&lt;span class="o"&gt;==&lt;/span&gt;&lt;span class="m"&gt;2&lt;/span&gt;.13.0 &lt;span class="kd"&gt;torchaudio&lt;/span&gt;&lt;span class="o"&gt;==&lt;/span&gt;&lt;span class="m"&gt;2&lt;/span&gt;.11.0 &lt;span class="kd"&gt;torchvision&lt;/span&gt;&lt;span class="o"&gt;==&lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;.28.0 &lt;span class="se"&gt;^
&lt;/span&gt;    &lt;span class="na"&gt;--index-url &lt;/span&gt;&lt;span class="kd"&gt;https&lt;/span&gt;://download.pytorch.org/whl/cu130
&lt;span class="kd"&gt;pip&lt;/span&gt; &lt;span class="kd"&gt;install&lt;/span&gt; &lt;span class="kd"&gt;triton&lt;/span&gt;&lt;span class="na"&gt;-windows&lt;/span&gt;&lt;span class="o"&gt;==&lt;/span&gt;&lt;span class="m"&gt;3&lt;/span&gt;.7.1.post27
&lt;span class="kd"&gt;pip&lt;/span&gt; &lt;span class="kd"&gt;install&lt;/span&gt; &lt;span class="kd"&gt;vllm&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;.27.1&lt;span class="na"&gt;-cp&lt;/span&gt;&lt;span class="m"&gt;313&lt;/span&gt;&lt;span class="na"&gt;-cp&lt;/span&gt;&lt;span class="m"&gt;313&lt;/span&gt;&lt;span class="na"&gt;-win&lt;/span&gt;_amd64.whl
&lt;span class="kd"&gt;pip&lt;/span&gt; &lt;span class="kd"&gt;install&lt;/span&gt; &lt;span class="kd"&gt;multi_turboquant&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;.1.0&lt;span class="na"&gt;-py&lt;/span&gt;&lt;span class="m"&gt;3&lt;/span&gt;&lt;span class="na"&gt;-none-any&lt;/span&gt;.whl
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both wheels are on the release page. Their published SHA-256 hashes are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;vllm-0.27.1-cp313-cp313-win_amd64.whl&lt;/code&gt;: &lt;code&gt;7c13ed44e94694478bdd4f5fcca23e2d66ba1e8fa9bccad9fddb8651d1b2447b&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;multi_turboquant-0.1.0-py3-none-any.whl&lt;/code&gt;: &lt;code&gt;5b310e05904b588539d9a8e3374dfa6c160f025f9c2099ba5c7877c79b2fa149&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your environment was built from older wheel metadata, there's an optional repair step: &lt;code&gt;pip install "llguidance&amp;gt;=1.7.0,&amp;lt;1.8.0" "xgrammar&amp;gt;=0.2.0,&amp;lt;1.0.0"&lt;/code&gt;. Windows reports its machine type as &lt;code&gt;AMD64&lt;/code&gt;, and older metadata skipped these packages because of that. The 0.27.1 wheel already has the correct markers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Using it
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;OpenAI-compatible HTTP.&lt;/strong&gt; The launcher (&lt;code&gt;vllm_launcher.py&lt;/code&gt;) serves &lt;code&gt;/v1/models&lt;/code&gt;, &lt;code&gt;/v1/chat/completions&lt;/code&gt; (streaming and non-streaming), &lt;code&gt;/v1/completions&lt;/code&gt;, &lt;code&gt;/health&lt;/code&gt;, and &lt;code&gt;/shutdown&lt;/code&gt;. Its default port is 8100 and it binds to &lt;code&gt;127.0.0.1&lt;/code&gt;. A quick test:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;
&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;http://127.0.0.1:8000/v1/chat/completions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;qwen3-14b&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;messages&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Explain CUDA streams in 3 sentences.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;max_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;choices&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It also parses tool calls from &lt;code&gt;&amp;lt;tool_call&amp;gt;&lt;/code&gt; tags (Qwen3 format) and from bare JSON, and returns them in &lt;code&gt;tool_calls&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Upstream &lt;code&gt;vllm serve&lt;/code&gt; works too. The docs still call the launcher the more reliable path on Windows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Python embedding.&lt;/strong&gt; The usage doc sets &lt;code&gt;VLLM_HOST_IP=127.0.0.1&lt;/code&gt;, adds the CUDA &lt;code&gt;bin&lt;/code&gt; and &lt;code&gt;torch\lib&lt;/code&gt; directories with &lt;code&gt;os.add_dll_directory&lt;/code&gt;, and then uses &lt;code&gt;LLM(...)&lt;/code&gt; / &lt;code&gt;generate()&lt;/code&gt; as usual.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Models.&lt;/strong&gt; Use a Hugging Face-format directory or repo ID with &lt;code&gt;config.json&lt;/code&gt; and Safetensors, AWQ, or GPTQ weights. The launcher rejects direct &lt;code&gt;.gguf&lt;/code&gt; files.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Memory tuning.&lt;/strong&gt; If the weights load but you get "No available memory for the cache blocks", the docs say to raise &lt;code&gt;--gpu-memory-utilization&lt;/code&gt; if VRAM is free (launcher default 0.6), or lower &lt;code&gt;--max-model-len&lt;/code&gt;, &lt;code&gt;--max-num-seqs&lt;/code&gt;, and &lt;code&gt;--max-num-batched-tokens&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RTX 20xx (Turing).&lt;/strong&gt; Add &lt;code&gt;--turing-compat&lt;/code&gt;. It selects &lt;code&gt;TRITON_ATTN&lt;/code&gt;, float16 KV cache, 32-token blocks, eager mode, 0.89 GPU utilization, one sequence, and 2,048 batched tokens. Any value you pass explicitly overrides the profile.&lt;/p&gt;

&lt;h2&gt;
  
  
  KV cache options
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;--kv-cache-dtype&lt;/code&gt; accepts &lt;code&gt;auto&lt;/code&gt;, the fp8 variants, four upstream TurboQuant variants (&lt;code&gt;turboquant_k8v4&lt;/code&gt;, &lt;code&gt;turboquant_4bit_nc&lt;/code&gt;, &lt;code&gt;turboquant_k3v4_nc&lt;/code&gt;, &lt;code&gt;turboquant_3bit_nc&lt;/code&gt;), and six methods from my &lt;a href="https://github.com/aivrar/multi-turboquant" rel="noopener noreferrer"&gt;Multi-TurboQuant&lt;/a&gt; library (&lt;code&gt;isoquant3/4&lt;/code&gt;, &lt;code&gt;planarquant3/4&lt;/code&gt;, &lt;code&gt;turboquant25/35&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;Be honest with yourself about the trade-off. The six Multi-TurboQuant methods doubled KV cache capacity in the README's RTX 3090 test (16,336 → 32,672 tokens). But their encode/decode still runs in PyTorch, and throughput drops roughly 30–300×. The upstream &lt;code&gt;turboquant_*&lt;/code&gt; variants use fused Triton kernels and don't pay that cost. For normal interactive use, stick with &lt;code&gt;auto&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;There's also experimental prompt-KV offload, which is off by default: &lt;code&gt;--kv-offload cpu-lru|cpu-arc|fs-lru|fs-arc&lt;/code&gt;. The filesystem modes need &lt;code&gt;--kv-offload-fs-root&lt;/code&gt;, and that directory has &lt;strong&gt;no automatic quota or cleanup&lt;/strong&gt;, so keep an eye on its size. Offload only helps when long prompt prefixes repeat. It doesn't speed up uncached generation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Single GPU per process.&lt;/strong&gt; For multiple GPUs, run separate instances on different ports and load-balance them yourself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No FlashInfer, FlashAttention 3/4, fastsafetensors, DeepGEMM, or NIXL.&lt;/strong&gt; None of them have Windows support here. FlashAttention 2 works.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blackwell FP8/NVFP4:&lt;/strong&gt; native Blackwell FP4 acceleration isn't included. Some mixed FP8/NVFP4 models fail at startup on unmodified 0.27.1. The &lt;a href="https://github.com/aivrar/vllm-windows-build/blob/master/docs/troubleshooting.md#blackwell-fp8-nvfp4" rel="noopener noreferrer"&gt;troubleshooting doc&lt;/a&gt; has a Marlin-based workaround (confirmed for one model), and the 0.29.0 builds include the Marlin fallback fixes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cold start:&lt;/strong&gt; the first inference that uses Triton kernels can take about 1–2 minutes to compile.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Limited testing.&lt;/strong&gt; Validation is targeted. Most of it ran on one RTX 3090, with an earlier release also run on an RTX 3060, plus single-model confirmations from issue reporters on an RTX 2080 Ti, an RTX 5090, and an RTX PRO 5000 Blackwell. The README's timings are focused correctness checks, not broad benchmarks, and the 0.29.0 Python 3.14 build doesn't claim any speed advantage.&lt;/li&gt;
&lt;li&gt;Building from source is possible (VS 2022 + CUDA Toolkit + a pagefile, no sccache), but it's heavy. The build records for each release are in &lt;code&gt;docs/&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/aivrar/vllm-windows-build" rel="noopener noreferrer"&gt;https://github.com/aivrar/vllm-windows-build&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;More of my tools: &lt;a href="https://github.com/aivrar" rel="noopener noreferrer"&gt;https://github.com/aivrar&lt;/a&gt;&lt;/p&gt;

</description>
      <category>vllm</category>
      <category>windows</category>
      <category>cuda</category>
      <category>llm</category>
    </item>
    <item>
      <title>Running a local AI agent on Windows with LM Studio, no Docker or admin rights</title>
      <dc:creator>aivrar</dc:creator>
      <pubDate>Fri, 09 Oct 2026 06:47:45 +0000</pubDate>
      <link>https://dev.to/aivrar/running-a-local-ai-agent-on-windows-with-lm-studio-no-docker-or-admin-rights-1pam</link>
      <guid>https://dev.to/aivrar/running-a-local-ai-agent-on-windows-with-lm-studio-no-docker-or-admin-rights-1pam</guid>
      <description>&lt;p&gt;I wanted an AI agent on Windows that could use tools (files, a shell, web search, a browser) and talk to a model running on my own GPU. Most agent frameworks expect Linux, Docker, or a system Python you're allowed to install into. A lot of Windows machines don't have any of those, and on plenty of them you can't get admin rights either.&lt;/p&gt;

&lt;p&gt;So I built &lt;strong&gt;Portable Hermes Agent&lt;/strong&gt;. It's a portable Windows build of &lt;a href="https://github.com/NousResearch/hermes-agent" rel="noopener noreferrer"&gt;NousResearch/hermes-agent&lt;/a&gt; (MIT) with a desktop GUI, an LM Studio panel, a permissions panel, and some extra tools added on top.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Disclosure:&lt;/strong&gt; I'm the author. I built most of the custom tools, GUI, and integrations with heavy AI-assisted coding (Claude Code, as the README credits). The core agent comes from Nous Research. This post sticks to what the repo and its release notes document.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you get
&lt;/h2&gt;

&lt;p&gt;From the README, the portable build adds:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A dark-themed Tkinter GUI with chat, a sidebar, and session management. The interface can switch at runtime between English, Traditional Chinese, and Simplified Chinese.&lt;/li&gt;
&lt;li&gt;100 tools in 20+ toolsets. That includes 10 LM Studio tools (load/unload models, search Hugging Face, tokenize, embed, direct chat, API key management), workflows, a tool maker, an NVIDIA GPU status tool, a model switcher, and Serper search.&lt;/li&gt;
&lt;li&gt;Everything hermes-agent already ships: web search, file operations, browser automation, code execution, delegation, memory, skills, messaging, Home Assistant, and more.&lt;/li&gt;
&lt;li&gt;A permissions panel for file, network, and system access.&lt;/li&gt;
&lt;li&gt;Optional extension modules for TTS, music, and image generation (ComfyUI).&lt;/li&gt;
&lt;li&gt;A "guided mode" that works with no model connected and answers from a built-in user guide.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The LLM provider can be OpenRouter (cloud), LM Studio (local), or any OpenAI-compatible endpoint.&lt;/p&gt;

&lt;h2&gt;
  
  
  Requirements
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Windows 10 or 11&lt;/li&gt;
&lt;li&gt;An internet connection for cloud models, &lt;strong&gt;or&lt;/strong&gt; an NVIDIA GPU with 8 GB+ VRAM for local models&lt;/li&gt;
&lt;li&gt;No admin rights, no system Python, no Docker&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Step 1: Download the right zip
&lt;/h2&gt;

&lt;p&gt;Go to the &lt;a href="https://github.com/aivrar/portable-hermes-agent/releases/latest" rel="noopener noreferrer"&gt;Releases page&lt;/a&gt; and download the file named &lt;code&gt;portable-hermes-agent-v*.zip&lt;/code&gt;. At the time of writing that's v1.4.10. Release notes include it, along with a PDF manual.&lt;/p&gt;

&lt;p&gt;The release notes say this more than once: use the &lt;strong&gt;named portable zip&lt;/strong&gt;, not GitHub's automatic "Source code" archives. Those aren't the verified release package.&lt;/p&gt;

&lt;p&gt;Extract it to a normal user folder, for example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;C:\Users\YourName\Portable-Hermes-Agent
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Don't put it in protected locations like &lt;code&gt;C:\Program Files&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: First launch
&lt;/h2&gt;

&lt;p&gt;Double-click:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight batchfile"&gt;&lt;code&gt;&lt;span class="kd"&gt;START&lt;/span&gt;.bat
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On first launch, &lt;code&gt;START.bat&lt;/code&gt; runs the portable setup. It downloads embedded Python, the dependencies, the LM Studio SDK, and Node.js tools into that folder only. If you'd rather run setup yourself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight batchfile"&gt;&lt;code&gt;&lt;span class="kd"&gt;install&lt;/span&gt;.bat
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;or from PowerShell:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;\scripts\install.ps1&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After that, these are the launchers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight batchfile"&gt;&lt;code&gt;&lt;span class="kd"&gt;START&lt;/span&gt;.bat           :: &lt;span class="kd"&gt;easiest&lt;/span&gt; &lt;span class="kd"&gt;GUI&lt;/span&gt; &lt;span class="kd"&gt;launch&lt;/span&gt;
&lt;span class="kd"&gt;hermes_gui&lt;/span&gt;.bat      :: &lt;span class="kd"&gt;GUI&lt;/span&gt; &lt;span class="nb"&gt;mode&lt;/span&gt;
&lt;span class="kd"&gt;hermes&lt;/span&gt;.bat          :: &lt;span class="kd"&gt;CLI&lt;/span&gt; &lt;span class="nb"&gt;mode&lt;/span&gt;
&lt;span class="kd"&gt;UPDATE&lt;/span&gt;.bat          :: &lt;span class="kd"&gt;safest&lt;/span&gt; &lt;span class="kd"&gt;one&lt;/span&gt;&lt;span class="na"&gt;-click &lt;/span&gt;&lt;span class="kd"&gt;update&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Step 3: Set up LM Studio as the local backend
&lt;/h2&gt;

&lt;p&gt;This part runs on your own GPU.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Install &lt;a href="https://lmstudio.ai" rel="noopener noreferrer"&gt;LM Studio&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Download a model in LM Studio's Discover tab. The manual uses "Qwen 2.5 7B" at Q4 or Q6 as its starter example.&lt;/li&gt;
&lt;li&gt;In LM Studio's Developer tab, click &lt;strong&gt;Start Server&lt;/strong&gt;. The manual gives the default port as 1234.&lt;/li&gt;
&lt;li&gt;In Hermes, open the LM Studio panel. The README calls it &lt;strong&gt;Tools &amp;gt; LM Studio&lt;/strong&gt;. The manual also lists it as &lt;strong&gt;View &amp;gt; LM Studio (Local Models)&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The manual describes the panel like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;an endpoint field (default &lt;code&gt;http://localhost:1234&lt;/code&gt;) with a Connect button&lt;/li&gt;
&lt;li&gt;a status indicator (green dot when connected)&lt;/li&gt;
&lt;li&gt;a model browser, a GPU selector, and a context length setting&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Load Model&lt;/strong&gt; / &lt;strong&gt;Cancel Load&lt;/strong&gt;, &lt;strong&gt;Unload&lt;/strong&gt;, and &lt;strong&gt;Use for Chat&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;Pick your model and GPU, set the context length, click &lt;strong&gt;Load Model&lt;/strong&gt;, then click &lt;strong&gt;Use for Chat&lt;/strong&gt;. Your messages now go to the local model.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Context length matters here.&lt;/strong&gt; The agent sends a lot of tool definitions, so small contexts don't work. As of v1.4.8, the panel's default context matches the core's minimum of 64K, and the panel warns you before loading a model whose maximum context is too small. (The bundled manual still says 32,768 in places. Trust the newer release note.) More context uses more VRAM, so if a model doesn't fit, try a smaller quantization or a smaller model.&lt;/p&gt;

&lt;p&gt;If your LM Studio server needs an API key, you can enter it in the same panel. The panel saves the endpoint and key to &lt;code&gt;.lmstudio_config&lt;/code&gt; in the active &lt;code&gt;HERMES_HOME&lt;/code&gt; directory, outside the application source. Both update paths keep that file, and release archives leave it out. To remove a saved key, clear the field and click &lt;strong&gt;Save&lt;/strong&gt;. Keep that file private.&lt;/p&gt;

&lt;p&gt;On multi-GPU machines, the manual says that choosing a GPU in the panel disables the other GPUs for that load, which forces the model onto one card.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No GPU?&lt;/strong&gt; Use &lt;strong&gt;File &amp;gt; API Key Setup &amp;gt; OpenRouter&lt;/strong&gt;, paste a key, and start chatting. You can switch between cloud and local later, from the sidebar or by asking the agent (for example "Switch to my local model").&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: Set permissions before you give it real work
&lt;/h2&gt;

&lt;p&gt;An agent that can run commands and edit files needs limits, so open &lt;strong&gt;Tools &amp;gt; Permissions&lt;/strong&gt; before doing anything real. The manual lists categories for file reading, file writing, file deletion, package installation, command execution, and network access. Each one goes from Level 0 (disabled) to Level 4 (system/admin).&lt;/p&gt;

&lt;p&gt;Defaults per the manual:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;File reading:&lt;/strong&gt; App + Home (Hermes folder plus your user folder)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;File writing / deletion:&lt;/strong&gt; App Only (inside the Hermes folder)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Package installation:&lt;/strong&gt; App Only (into the bundled portable Python)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Command execution:&lt;/strong&gt; App + Safe (runs commands, no admin/system changes)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Network:&lt;/strong&gt; Web + APIs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you want to keep everything on your machine, Network Level 1 ("Local Only") limits it to localhost services like LM Studio and the extensions. I'd start with the defaults and raise one category at a time when a task actually needs it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 5 (optional): Extensions
&lt;/h2&gt;

&lt;p&gt;Three extensions are separate portable servers I also maintain:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Extension&lt;/th&gt;
&lt;th&gt;Port&lt;/th&gt;
&lt;th&gt;GPU&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/aivrar/portable-tts-server" rel="noopener noreferrer"&gt;TTS Server&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;8200&lt;/td&gt;
&lt;td&gt;4 GB+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/aivrar/portable-music-server" rel="noopener noreferrer"&gt;Music Server&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;9150&lt;/td&gt;
&lt;td&gt;4 GB+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://github.com/aivrar/comfyui-portable-installer" rel="noopener noreferrer"&gt;ComfyUI&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;5000&lt;/td&gt;
&lt;td&gt;6 GB+&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each one installs itself on first use into the &lt;code&gt;extensions/&lt;/code&gt; folder. You manage them from &lt;strong&gt;Tools &amp;gt; Extensions&lt;/strong&gt; (install, start/stop, status). The manual notes that the downloads are several GB. Remember they share the GPU with your LM Studio model, so you may need to unload one to run the other.&lt;/p&gt;

&lt;h2&gt;
  
  
  Workflows and the tool maker
&lt;/h2&gt;

&lt;p&gt;Two features beyond plain chat:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Workflow engine:&lt;/strong&gt; chain tool calls into pipelines with data flow, conditions, loops, parallel execution, error handling, and cron scheduling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool maker:&lt;/strong&gt; create new tools at runtime, either by wrapping a REST API or by writing a Python handler. They persist across sessions and reload automatically.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Updating
&lt;/h2&gt;

&lt;p&gt;There are two separate update channels, and updating one doesn't update the other:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;The portable distribution&lt;/strong&gt; (launchers, GUI, integrations, portable tools): close Hermes and run &lt;code&gt;UPDATE.bat&lt;/code&gt;, or &lt;code&gt;hermes.bat update --backup --yes&lt;/code&gt;. This preserves &lt;code&gt;.hermes/&lt;/code&gt;, custom tools, extensions, and &lt;code&gt;python_embedded/&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Upstream hermes-agent core:&lt;/strong&gt; start Hermes and ask it, e.g. "Check whether upstream Hermes Agent has updates. Do not install anything yet." Restart after a successful upstream update.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The README suggests running the portable updater when a release or security notice comes out (or every week or two). Recent releases (v1.4.9, v1.4.10) were mostly dependency security updates.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limitations
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Windows only&lt;/strong&gt;, and local models need an &lt;strong&gt;NVIDIA GPU&lt;/strong&gt; (8 GB+ suggested).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context is the main constraint locally.&lt;/strong&gt; The tool definitions are large, so you need a model and GPU that can handle about 64K context. That rules out a lot of small-VRAM setups.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Authenticated LM Studio model loading through the SDK&lt;/strong&gt; needs an LM Studio Python SDK version that supports &lt;code&gt;api_token&lt;/code&gt;. The v1.4.8 notes say the SDK installed at the time didn't, and that live auth enforcement wasn't tested. REST model discovery and chat still work without the SDK.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The GUI is Tkinter.&lt;/strong&gt; It works, but it's plain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Local model quality varies a lot.&lt;/strong&gt; The manual itself notes quality differs a lot between models. Expect smaller local models to be less dependable at multi-step tool use.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;p&gt;Repo: &lt;a href="https://github.com/aivrar/portable-hermes-agent" rel="noopener noreferrer"&gt;https://github.com/aivrar/portable-hermes-agent&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;More of my tools: &lt;a href="https://github.com/aivrar" rel="noopener noreferrer"&gt;https://github.com/aivrar&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>windows</category>
      <category>llm</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
