DEV Community

Cover image for vLLM GGUF FAQ: Ten Search Questions, Answered
Mr Say Nothing
Mr Say Nothing

Posted on Originally published at mrsaynothing.dev

vLLM GGUF FAQ: Ten Search Questions, Answered

These ten questions are not invented — every one is a real query from this site's own search console, and for a long stretch we ranked for each with zero clicks. The answers come from the official vLLM docs and our tested Can vLLM Run GGUF? Yes — on GPU Only deep dive.

1. Can vLLM run GGUF?

Yes — through the official vllm-gguf-plugin, on NVIDIA and AMD GPUs. Both halves matter: the plugin is vLLM's own, and the hardware table stops at GPU.

2. Does vLLM support GGUF on CPU?

No. No CPU kernel sits behind the GGUF path — x86/Arm CPUs and Intel GPUs are outside the supported table. This is the trap behind the classic report: the same GGUF that fails in vLLM on a laptop runs fine in Ollama, because GGUF's home turf is CPU-first engines. The file was never the problem.

3. How do you serve a GGUF model with vLLM?

vllm serve unsloth/Qwen3-0.6B-GGUF:Q4_K_M --tokenizer Qwen/Qwen3-0.6B
Enter fullscreen mode Exit fullscreen mode

The syntax is repo:quant_type. The trap is --tokenizer: it names the original model repo, never the GGUF repo — GGUF files carry weights, and vLLM wants the real tokenizer from the source model.

4. Which GGUF quants does vLLM support?

The K-quants people actually download — Q4_K_M and friends — work; exotic schemes may not. llama.cpp stays the reference implementation, so new quant schemes land there first and reach the plugin later.

5. Is vLLM's GGUF support production-ready?

The docs answer this one themselves: "GGUF support in vLLM is highly experimental and under-optimized." That sentence sits in the first paragraph of the official GGUF page — more honesty than most experimental features get. Treat the plugin as a way to reuse files you already have, and keep the exit to llama.cpp open.

6. Does vLLM load GGUF faster than llama.cpp?

No — and speed was never the point. llama.cpp memory-maps the file, so a GGUF starts lazily; vLLM loads it like any other checkpoint. GGUF inside vLLM is a VRAM-footprint feature, not a throughput play.

7. vLLM or llama.cpp for GGUF — which wins?

Same file, opposite instincts: llama.cpp treats GGUF as the format, vLLM treats it as an option. One GPU serving several people → vLLM plus the plugin. One person, any hardware, widest quant choice → llama.cpp, one binary.

8. Why does my GGUF file fail in vLLM?

Three real causes: plain-path serve syntax instead of repo:quant_type, a missing or wrong --tokenizer, or hardware outside the supported table. Rarer fourth: a quant scheme the plugin doesn't cover.

9. Does Ollama run GGUF?

Yes — GGUF is Ollama's native format. The same file that fails in vLLM on a CPU box usually just works there.

10. Can several users share one vLLM GGUF server?

That is the actual reason to do this: continuous batching lets one GPU serve concurrent users off the same GGUF — the users queue instead of the model reloading.

The full head-to-head — CPU, concurrency, quant coverage, setup — plus the failure-mode table live on the canonical post. One post a day, receipts included: the letters.

Top comments (2)

Collapse
 
mrsaynothing profile image
Mr Say Nothing •

All ten questions came from this site's own search console — every one a query the site already ranked for with zero clicks. The surprise doing the editing: the mandatory --tokenizer flag trips more people than CPU support does. Which of the two bit you?

Some comments may only be visible to logged-in visitors. Sign in to view all comments.