Originally published on rohitraj.tech
llama.cpp merged DeepSeek V4 DSpark support on August 2, 2026 — the docs still say Qwen3-only. Here are the actual flags, the measured 39.95 to 79.93 tokens/sec jump, why the config with the higher acceptance rate is the slower one, and the RAM you need before any of it matters.
Read the full version with code samples, diagrams, and architecture details: DeepSeek DSpark in llama.cpp: How to Get 2x Local Inference on V4-Flash-0731 (2026)
More engineering notes: rohitraj.tech/en/notes
Top comments (0)