Spark-X2.5-4B is a 4B open model that goes toe-to-toe with models 2–3× its size on agents, coding, and math. Here's the full benchmark table vs comparable open models:
🤖 Agents
The 4B leads its size class and beats much larger models:
- τ³-bench 30.4 (Qwen3.5-9B 9.3, Gemma4-12B 13.3)
- BrowseComp 40.9 (8.3 / 10.0)
- MCP-Atlas 54.6 · VitaBench2.0 25.2 · Workspace Bench 31.2
Deeply integrated with popular agent harnesses — Codex, Claude Code, OpenClaw, and Hermes.
💻 Code (at just 4B)
- SWE-Bench Pro 44.4 — ahead of Qwen3.5-9B (33.8) and Gemma4-12B (21.9)
- SWE-Bench Multilingual 53.3 — tops the group
- SWE-Bench Verified 41.6
🧮 Math
- AIME 2026 90.7
- HMMT Feb 2026 81.2
- IMO-AnswerBench 74.2
Each ahead of models 2–3× larger. The 1.7B also holds its own against 2B-class peers.
Get it
- GitHub: https://github.com/XHToken/Spark-X2.5
- Hugging Face: https://huggingface.co/collections/XHToken/spark-x25
- API on MaaS: https://maas.xfyun.cn/modelSquare
Figures from the official Spark-X2.5 release. Higher is better. ` denotes results reported from publicly-released model cards / papers; -` denotes scores not yet available; all evaluations run in thinking mode.*

Top comments (0)