Can a 1.7B model really run end to end?
This short recaps a public participant case: on a Colab NVIDIA L4, the developer built the XHToken llama.cpp fork with CUDA, ran Spark-X2.5-1.7B (BF16 GGUF), and documented the real prompt, output, speed observations, memory use, and limitations.
The point isn't a hero number — it's a reproducible case: run it for real, document it honestly, and share what the next developer can actually rebuild.
Join HER Hack-Astron #5
Turn a real inference, agent, coding, long-context, multilingual, quantization, training, or safety experiment into a reproducible case.
- Publish your case as a Discussion — see the example: Case Discussion #1
- Submit it: publish your Discussion first, then reply on the challenge issue with its direct link — Challenge Issue #3
⏰ Deadline: September 6, 2026, 24:00 (Beijing Time).
Run it for real, document it honestly, and share what the next developer can reproduce.
Spark-X2.5 — GitHub: https://github.com/XHToken/Spark-X2.5 · Hugging Face: https://huggingface.co/xhtoken
Top comments (0)