<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alfred Odong</title>
    <description>The latest articles on DEV Community by Alfred Odong (@alfred_odong_322108a5cc3d).</description>
    <link>https://dev.to/alfred_odong_322108a5cc3d</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4170214%2F26201fda-6ace-4701-a0b1-b6dafb8d2da7.png</url>
      <title>DEV Community: Alfred Odong</title>
      <link>https://dev.to/alfred_odong_322108a5cc3d</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/alfred_odong_322108a5cc3d"/>
    <language>en</language>
    <item>
      <title>I put seven LLMs on a USB stick. Here is everything that broke.</title>
      <dc:creator>Alfred Odong</dc:creator>
      <pubDate>Thu, 08 Oct 2026 06:09:51 +0000</pubDate>
      <link>https://dev.to/alfred_odong_322108a5cc3d/i-put-seven-llms-on-a-usb-stick-here-is-everything-that-broke-39cc</link>
      <guid>https://dev.to/alfred_odong_322108a5cc3d/i-put-seven-llms-on-a-usb-stick-here-is-everything-that-broke-39cc</guid>
      <description>&lt;p&gt;The idea fits in one sentence: plug a stick into any computer, double-click one&lt;br&gt;
file, and a real language model runs on that machine. No account, no internet,&lt;br&gt;
no install, and nothing left behind when you unplug it.&lt;/p&gt;

&lt;p&gt;That sentence took weeks to make true. Not because the AI part is hard, since&lt;br&gt;
llamafile and llama.cpp do the heavy lifting. It took weeks because every&lt;br&gt;
operating system has its own way of quietly refusing to run a program off a&lt;br&gt;
USB stick. Most of those refusals look like something else entirely.&lt;/p&gt;

&lt;p&gt;This is the list. Every item happened on real hardware, and the fix for each&lt;br&gt;
is in the kit.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. The server that exits successfully and does nothing
&lt;/h2&gt;

&lt;p&gt;On Windows, &lt;code&gt;llama-server.exe&lt;/code&gt; started, printed nothing, opened no window, and&lt;br&gt;
exited with &lt;strong&gt;code 0&lt;/strong&gt;. Code 0 means success. It looked exactly like a&lt;br&gt;
corrupted download, so I re-downloaded it. Twice.&lt;/p&gt;

&lt;p&gt;The real cause was three missing Microsoft runtime DLLs: &lt;code&gt;MSVCP140.dll&lt;/code&gt;,&lt;br&gt;
&lt;code&gt;VCRUNTIME140.dll&lt;/code&gt; and &lt;code&gt;VCRUNTIME140_1.dll&lt;/code&gt;. Any machine that has ever&lt;br&gt;
installed a game or Visual Studio has them. A clean machine doesn't, and the&lt;br&gt;
loader fails before the program can even report an error.&lt;/p&gt;

&lt;p&gt;The fix is to ship the DLLs next to the &lt;code&gt;.exe&lt;/code&gt;. I didn't just assume that&lt;br&gt;
works: I uninstalled the redistributable in a test VM, confirmed the registry&lt;br&gt;
key and the System32 copies were gone, and ran it again. It worked. (Those DLLs&lt;br&gt;
also have to be redistributed under Microsoft's actual terms, not copied out of&lt;br&gt;
a System32 folder. That cost its own afternoon.)&lt;/p&gt;

&lt;h2&gt;
  
  
  2. The engine that can't start other programs on Windows
&lt;/h2&gt;

&lt;p&gt;llamafile is a wonderful trick: one file that runs on macOS, Linux and&lt;br&gt;
Windows. On Windows its file tools worked fine. But every shell command the&lt;br&gt;
model tried to run returned &lt;code&gt;exit -1&lt;/code&gt;, while the same command typed by hand&lt;br&gt;
worked perfectly.&lt;/p&gt;

&lt;p&gt;Windows has no &lt;code&gt;fork()&lt;/code&gt;. llamafile's portability layer emulates it, and&lt;br&gt;
starting a child process is exactly where that emulation breaks. So the drive&lt;br&gt;
carries a second engine just for Windows, llama.cpp's own&lt;br&gt;
&lt;code&gt;llama-server.exe&lt;/code&gt;, and the Windows launchers prefer it.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Linux has no double-click, and GNOME closed the last door
&lt;/h2&gt;

&lt;p&gt;macOS runs a &lt;code&gt;.command&lt;/code&gt; file on double-click, and Windows runs a &lt;code&gt;.bat&lt;/code&gt;.&lt;br&gt;
Linux never had an equivalent for &lt;code&gt;.sh&lt;/code&gt; scripts. And on Ubuntu 26.04, GNOME's&lt;br&gt;
file manager won't launch a &lt;code&gt;.desktop&lt;/code&gt; file &lt;em&gt;or&lt;/em&gt; an executable script from&lt;br&gt;
anywhere outside the standard app folders. No setting, no "Allow Launching",&lt;br&gt;
no trust flag brings it back. Both launchers on the drive open in a text&lt;br&gt;
editor.&lt;/p&gt;

&lt;p&gt;There's no fix from the drive's side. The kit says so plainly and gives you&lt;br&gt;
two ways around it: a one-line terminal command, or a one-time installer that&lt;br&gt;
puts a launcher where GNOME will run it.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. FAT32 makes exactly one file executable, and it's the Windows one
&lt;/h2&gt;

&lt;p&gt;When a Linux desktop mounts a FAT32 stick, it marks only files ending in&lt;br&gt;
&lt;code&gt;.exe&lt;/code&gt;, &lt;code&gt;.com&lt;/code&gt; or &lt;code&gt;.bat&lt;/code&gt; as executable. On a FAT32 copy of the drive, that&lt;br&gt;
makes &lt;code&gt;WINDOWS-start-ai.bat&lt;/code&gt; the only runnable file on the stick, and running&lt;br&gt;
the Linux launcher fails with "Permission denied".&lt;/p&gt;

&lt;p&gt;exFAT doesn't do this. It also lifts FAT32's 4 GB limit on a single file,&lt;br&gt;
which matters once a model is larger than that. Format the stick as exFAT.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The copy that wasn't on the stick at all
&lt;/h2&gt;

&lt;p&gt;I copied the kit, checked it, and everything was there: right sizes, right&lt;br&gt;
checksums. I ejected, plugged it back in, and found a drive full of &lt;strong&gt;0-byte&lt;br&gt;
files&lt;/strong&gt; that still had every filename correct.&lt;/p&gt;

&lt;p&gt;The check had read the files back from the operating system's write cache,&lt;br&gt;
not from the stick. It happened twice in one afternoon. The rule now: eject,&lt;br&gt;
replug, &lt;em&gt;then&lt;/em&gt; verify.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. The launcher that hangs forever on a computer with no browser set
&lt;/h2&gt;

&lt;p&gt;The Windows launcher opened the chat page with &lt;code&gt;start "" http://...&lt;/code&gt;. On a&lt;br&gt;
machine with no default browser, that call never returns, and the whole&lt;br&gt;
launcher freezes with no error.&lt;/p&gt;

&lt;p&gt;The fix is to hand the browser launch off through PowerShell, and to always&lt;br&gt;
print the address so you can open it yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. A cheap stick is a speed limit, and it flips a rule I believed
&lt;/h2&gt;

&lt;p&gt;My test stick reads at &lt;strong&gt;34 MB/s&lt;/strong&gt;. The whole model is read off the stick&lt;br&gt;
every time it starts, so a 2.6 GB model takes over a minute before the first&lt;br&gt;
answer. A good USB 3 drive does it in seconds.&lt;/p&gt;

&lt;p&gt;On my home server, a bigger "mixture of experts" model was actually &lt;em&gt;faster&lt;/em&gt;,&lt;br&gt;
because memory was the bottleneck there. On a USB stick, file size is the only&lt;br&gt;
thing that matters: an 18 GB model would need about nine minutes just to load.&lt;br&gt;
So the kit ships small dense models, and the advice is to upgrade the drive&lt;br&gt;
before you upgrade the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. The model that thought itself out of a job
&lt;/h2&gt;

&lt;p&gt;Qwen3 is a reasoning model, so it thinks before every answer. Asked to open&lt;br&gt;
the Recycle Bin, it spent &lt;strong&gt;441 tokens&lt;/strong&gt; deliberating and still got it wrong.&lt;br&gt;
With reasoning off, the same kind of question took about a dozen tokens.&lt;/p&gt;

&lt;p&gt;On a small model with a limited context window, that's the difference between&lt;br&gt;
an assistant and something that fills its memory before it answers. Every&lt;br&gt;
launcher turns reasoning off. A related default: llama-server splits its&lt;br&gt;
context between four parallel users. One person on a stick needs one, so the&lt;br&gt;
launchers set it to one and get four times the room.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. The macOS crash that came back as a zombie
&lt;/h2&gt;

&lt;p&gt;On macOS, llamafile sometimes crashes inside &lt;code&gt;fork()&lt;/code&gt; when the model runs a&lt;br&gt;
shell command. That's an upstream bug I can't fix, so the Mac launcher&lt;br&gt;
restarts the server automatically. I tested it by killing the server myself,&lt;br&gt;
and it came back on the same port within two seconds.&lt;/p&gt;

&lt;p&gt;The bug has a second form. Sometimes the forked copy doesn't crash but&lt;br&gt;
deadlocks, spinning at 100% CPU and ignoring the normal stop signal. Two of&lt;br&gt;
them ran for 8.5 hours, held the port, and pushed the next launch onto a&lt;br&gt;
different port without saying so. The launcher now finds these by their&lt;br&gt;
fingerprint (the right port, no parent process, exactly one thread; a live&lt;br&gt;
server has about 17) and cleans them up.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. The small things that each cost an hour
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Duplicate lines in the memory file&lt;/strong&gt; sent the model into an endless loop.
Its edit tool needs a unique line to anchor on, so the model kept retrying
the same failing edit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A third &lt;code&gt;-------&lt;/code&gt; divider&lt;/strong&gt; in the system prompt file silently cut off
everything after it, including the whole table of command recipes. The
build script now refuses any prompt file that doesn't have exactly two.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Windows PowerShell 5.1 reads UTF-8 files as ANSI&lt;/strong&gt; unless told otherwise,
so an em dash came out as &lt;code&gt;√¢?"&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;My own test harness capped the model at 5 steps.&lt;/strong&gt; I spent a while
calling that the model's "multistep ceiling". With 10 steps, the same model
went from 3/5 to 5/5.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  11. The speech-to-text engine that refuses every 2048th sample
&lt;/h2&gt;

&lt;p&gt;Version 1.1 adds hearing: a whisper.cpp build (whisperfile) that turns voice&lt;br&gt;
memos and meetings into text, and a push-to-talk page that runs the whole&lt;br&gt;
loop locally. Its decoder fails with &lt;code&gt;failed to read pcm frames: At end&lt;/code&gt; on&lt;br&gt;
some files and not others, with no pattern I could see.&lt;/p&gt;

&lt;p&gt;The pattern, after a sweep of 40 generated files: it fails whenever the&lt;br&gt;
sample count is an exact multiple of &lt;strong&gt;2048&lt;/strong&gt;. A browser's audio capture&lt;br&gt;
buffer is 4096 samples, so every recording from the voice page hit it. It&lt;br&gt;
also fails on WAVs that carry a metadata chunk, which ffmpeg adds when the&lt;br&gt;
source is an &lt;code&gt;.m4a&lt;/code&gt;. The verb now re-encodes everything, strips metadata, and&lt;br&gt;
pads one extra sample if the count lands on the wrong number. On a Mac with&lt;br&gt;
no ffmpeg it falls back to &lt;code&gt;afconvert&lt;/code&gt; and rewrites the 68-byte header it&lt;br&gt;
produces into the plain 44-byte one the decoder expects.&lt;/p&gt;

&lt;p&gt;Two more from the same release: &lt;code&gt;spd-say -w&lt;/code&gt; on a headless Linux box blocks&lt;br&gt;
forever, so the talk-back verb now gives it a time budget; and a server&lt;br&gt;
started from a tool that runs at &lt;code&gt;nice 5&lt;/code&gt; gets its child processes parked on&lt;br&gt;
efficiency cores under load, which made a 3-second transcription take 37.&lt;/p&gt;




&lt;h2&gt;
  
  
  What it is now
&lt;/h2&gt;

&lt;p&gt;Seven models with a picker at launch, from Qwen3-1.7B (fast, runs on&lt;br&gt;
anything) to Qwen3-8B (smartest, wants about 16 GB of RAM), plus a Vision&lt;br&gt;
model that reads the photos, receipts and screenshots you attach. It&lt;br&gt;
transcribes audio and reads answers aloud, and a voice page lets you hold a&lt;br&gt;
button, ask, and hear the reply, with the microphone audio going to a server&lt;br&gt;
on localhost and nowhere else. Launchers for macOS, Windows and Linux, a&lt;br&gt;
files-only mode, and an agent mode that asks your permission in the browser&lt;br&gt;
before &lt;strong&gt;every&lt;/strong&gt; shell command it runs. A memory file lives on the drive, so&lt;br&gt;
the assistant remembers you from machine to machine.&lt;/p&gt;

&lt;p&gt;It was tested on an Apple Silicon Mac, Ubuntu 26.04, and Windows Server 2025&lt;br&gt;
running off the physical stick. The models are small: they're useful for&lt;br&gt;
writing, summarising, editing files, reading a receipt and running a computer&lt;br&gt;
from plain English, but they're not frontier models and they'll sometimes be&lt;br&gt;
confidently wrong. Nothing on a USB stick is a frontier model.&lt;/p&gt;

&lt;p&gt;If you'd rather build it yourself, the build kit downloads every model and&lt;br&gt;
engine straight from the people who publish them and checks each against the&lt;br&gt;
publisher's SHA-256. If you just want it working, there's a ready-made 13 GB&lt;br&gt;
drive image you copy onto a stick.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Portable AI on a Drive, $59:&lt;/strong&gt; &lt;a href="https://primeagent2.gumroad.com/l/objkjr" rel="noopener noreferrer"&gt;https://primeagent2.gumroad.com/l/objkjr&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Launch code &lt;strong&gt;USB44&lt;/strong&gt; takes $15 off until 27 October.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://primeagent2.gumroad.com/p/i-put-seven-llms-on-a-usb-stick-here-is-everything-that-broke" rel="noopener noreferrer"&gt;Gumroad&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>ai</category>
      <category>linux</category>
      <category>windows</category>
    </item>
  </channel>
</rss>
