DEV Community

SparkLLM
SparkLLM

Posted on

Spark-X2.5-4B: a 4B model matching 2–3 larger models on agents, code & math

Spark-X2.5-4B is a 4B open model that goes toe-to-toe with models 2–3× its size on agents, coding, and math. Here's the full benchmark table vs comparable open models:

Spark-X2.5 benchmarks across agent, code, math and general tasks

🤖 Agents

The 4B leads its size class and beats much larger models:

  • τ³-bench 30.4 (Qwen3.5-9B 9.3, Gemma4-12B 13.3)
  • BrowseComp 40.9 (8.3 / 10.0)
  • MCP-Atlas 54.6 · VitaBench2.0 25.2 · Workspace Bench 31.2

Deeply integrated with popular agent harnesses — Codex, Claude Code, OpenClaw, and Hermes.

💻 Code (at just 4B)

  • SWE-Bench Pro 44.4 — ahead of Qwen3.5-9B (33.8) and Gemma4-12B (21.9)
  • SWE-Bench Multilingual 53.3 — tops the group
  • SWE-Bench Verified 41.6

🧮 Math

  • AIME 2026 90.7
  • HMMT Feb 2026 81.2
  • IMO-AnswerBench 74.2

Each ahead of models 2–3× larger. The 1.7B also holds its own against 2B-class peers.

Get it

Figures from the official Spark-X2.5 release. Higher is better. ` denotes results reported from publicly-released model cards / papers; -` denotes scores not yet available; all evaluations run in thinking mode.*

Top comments (0)