Official llama.cpp support for Spark-X2.5 is here — so you can run the models locally, no cloud required.
This short walks through a real Windows CPU demo using the community 1.7B Q4_K_M GGUF and official build b10833:
- Download the quantized GGUF
- Launch a local server with llama.cpp
- Ask a question — running entirely on CPU
Both sizes are available to explore — the 4B and the 1.7B.
Get started
- 📦 Model: https://github.com/XHToken/Spark-X2.5
- ⚙️ Runtime (llama.cpp): https://github.com/ggml-org/llama.cpp
Try it locally and share your feedback, experiences, and creations — we'd love to see what you build.
Top comments (0)