These ten questions are not invented — every one is a real query from this site's own search console, and for a long stretch we ranked for each with zero clicks. The answers come from the official vLLM docs and our tested Can vLLM Run GGUF? Yes — on GPU Only deep dive.
1. Can vLLM run GGUF?
Yes — through the official vllm-gguf-plugin, on NVIDIA and AMD GPUs. Both halves matter: the plugin is vLLM's own, and the hardware table stops at GPU.
2. Does vLLM support GGUF on CPU?
No. No CPU kernel sits behind the GGUF path — x86/Arm CPUs and Intel GPUs are outside the supported table. This is the trap behind the classic report: the same GGUF that fails in vLLM on a laptop runs fine in Ollama, because GGUF's home turf is CPU-first engines. The file was never the problem.
3. How do you serve a GGUF model with vLLM?
vllm serve unsloth/Qwen3-0.6B-GGUF:Q4_K_M --tokenizer Qwen/Qwen3-0.6B
The syntax is repo:quant_type. The trap is --tokenizer: it names the original model repo, never the GGUF repo — GGUF files carry weights, and vLLM wants the real tokenizer from the source model.
4. Which GGUF quants does vLLM support?
The K-quants people actually download — Q4_K_M and friends — work; exotic schemes may not. llama.cpp stays the reference implementation, so new quant schemes land there first and reach the plugin later.
5. Is vLLM's GGUF support production-ready?
The docs answer this one themselves: "GGUF support in vLLM is highly experimental and under-optimized." That sentence sits in the first paragraph of the official GGUF page — more honesty than most experimental features get. Treat the plugin as a way to reuse files you already have, and keep the exit to llama.cpp open.
6. Does vLLM load GGUF faster than llama.cpp?
No — and speed was never the point. llama.cpp memory-maps the file, so a GGUF starts lazily; vLLM loads it like any other checkpoint. GGUF inside vLLM is a VRAM-footprint feature, not a throughput play.
7. vLLM or llama.cpp for GGUF — which wins?
Same file, opposite instincts: llama.cpp treats GGUF as the format, vLLM treats it as an option. One GPU serving several people → vLLM plus the plugin. One person, any hardware, widest quant choice → llama.cpp, one binary.
8. Why does my GGUF file fail in vLLM?
Three real causes: plain-path serve syntax instead of repo:quant_type, a missing or wrong --tokenizer, or hardware outside the supported table. Rarer fourth: a quant scheme the plugin doesn't cover.
9. Does Ollama run GGUF?
Yes — GGUF is Ollama's native format. The same file that fails in vLLM on a CPU box usually just works there.
10. Can several users share one vLLM GGUF server?
That is the actual reason to do this: continuous batching lets one GPU serve concurrent users off the same GGUF — the users queue instead of the model reloading.
The full head-to-head — CPU, concurrency, quant coverage, setup — plus the failure-mode table live on the canonical post. One post a day, receipts included: the letters.
Top comments (2)
All ten questions came from this site's own search console — every one a query the site already ranked for with zero clicks. The surprise doing the editing: the mandatory --tokenizer flag trips more people than CPU support does. Which of the two bit you?
Some comments may only be visible to logged-in visitors. Sign in to view all comments.