Running LLMs locally on Linux: what actually works on a Raspberry Pi
A 5B-parameter model can run on a Raspberry Pi 5 every day. It is not fast. It is also completely offline, costs nothing per token, and never phones home. That trade is worth making for a specific class of work, and worthless for everything else. Below is what the WIAIA community has found after several months of experimenting.
The hardware reality nobody mentions
Before any software, the constraint. A Pi 5 has 8GB of RAM and shares it with the rest of the system. When one loads a 4-bit quantized 5B model, the weights alone eat roughly 3.5 to 4GB. Add context KV cache and space is already tight. A usable context window lands around 2 to 3k tokens before generation starts failing and the process gets OOM-killed.
This is the single most important fact about local inference on a Pi, and it is not about the GPU or the quantization method. It is that roughly 4GB is available, and everything competes for it.
For actual throughput, stop chasing tokens per second. Raw speed on this hardware is poor. What matters more is that a request never leaves the machine, so a job can be left running and the result checked hours later.
What people actually run
Three tools have earned their place:
Ollama is the common default. Model management is two commands, it has a stable HTTP API on port 11434, and it runs behind systemd so it survives reboots. For anything scripted, this is the one worth starting with. It is also the one that needs care for sensitive work, because it can bind to all interfaces depending on the install. Check ollama.allowed_origins in the config.
llama.cpp via llama-server is what to reach for when control over sampling matters. The GGUF file is yours, every sampling parameter is exposed as a flag, and threads can be pinned to specific cores. It sits closer to the metal and closer to what perplexity actually means.
For anything expecting an OpenAI-compatible endpoint, both of the above expose one. Point a client at http://localhost:11434/v1 or the llama.cpp equivalent and the difference stops mattering. Most tooling never notices.
Quantization: pick 4-bit and move on
Q4_K_M is the pragmatic default. It is the sweet spot where the model barely degrades from the original and the file size stays manageable. For a 5B model that lands around 3.5GB on disk, which is what makes it fit on a Pi at all.
Going to Q8 doubles the size with no perceptible difference on typical workloads. Going to Q2 saves space but output degrades into word salad. There is no reason to be a martyr about it. Q4_K_M, load it, move on.
The thing worth knowing is that quantization is not free. Information is lost, and on reasoning-heavy tasks the loss surfaces as confidently wrong answers rather than obvious garbage. A badly quantized model does not say "I don't know." It says something plausible in the same register as the rest of its output.
Prompting a small model is a different job
This takes the longest to accept. A 5B model does not follow complex instructions. Multi-part prompts with formatting requirements produce output that satisfies roughly one constraint and quietly ignores the rest.
What works:
- One task per prompt. Not "summarize this and list the action items and format as markdown." Just "summarize this."
- Short output targets. "Three sentences" gets obeyed. "A thorough analysis" produces rambling until context runs out and the output gets truncated mid-sentence.
- Output format in the system prompt and nowhere else. Redundancy in a small model reads as emphasis, not confirmation.
- A one-shot example when format matters. One. Not three.
- Stop sequences instead of instructions to stop. Ending on a newline and setting
stop=["\n\n\n"]beats asking for concise output in prose.
Where the Pi setup beats a hosted API
This is the underrated part.
Data stays put. A model can be run over logs, config files, and half-written notes containing credentials and internal hostnames. None of it goes to anyone. That is the whole reason to buy the hardware.
No rate limits, no per-token cost, no quota. Jobs can queue up overnight without a single throttle error.
It works on a dead network. Local setups keep going when connectivity does not.
It is auditable. The weights are a file on disk. What is running is known, and it can be hashed.
Where it loses, badly
Local inference on Pi hardware is not competitive for anything latency-sensitive. Interactive autocomplete in an editor is out. Code completion across a large file is out. Anything where a human is waiting is out.
Suited workloads: summarizing a 40-page document, classifying a few thousand files, drafting from notes already written, batch generating variations for review. Hours of unattended work where a 1.5 tokens per second generation rate is fine.
Quality is capped. A 5B model will lose to a frontier model on anything requiring nuance or long reasoning. Worth accepting rather than pretending otherwise. If an answer needs to be right about something subtle, a bigger model is the right tool.
Setup, to try it
Install Ollama, pull a model under 6B parameters, and set the context window explicitly rather than accepting the default:
curl -fsSL https://ollama.com/install.sh | sh
ollama pull qwen2.5:3b
OLLAMA_CONTEXT_LENGTH=2048 systemctl restart ollama
Set OLLAMA_NUM_PARALLEL=1. On 8GB shared memory, parallel requests multiply KV cache usage and hit the OOM killer faster than expected.
Then test before trusting it:
time curl http://localhost:11434/api/generate -d '{
"model": "qwen2.5:3b",
"prompt": "Summarize in three sentences: the tradeoff of running models locally.",
"stream": false
}'
If that round trips without a kernel OOM message in dmesg, the setup is working.
The bigger picture
An experimental setup like this was expected to migrate everything to a hosted model once the wait times became annoying. What actually happened is that the work split cleanly in two. Hosted APIs took everything interactive and judgment-heavy. The Pi took the volume work that is slow, repetitive, and touches data better kept local.
That split is the real lesson. Local inference is not a cheaper version of the cloud. It is a different tool for a different job, and the useful decision is which jobs go where.
Interested in contributing to WIAIA or sharing what has worked in your own setup? Visit wiaia.github.io to learn more about the community, the mentorship program, and how to get involved.
Top comments (0)