A correct mathematical result can still fail a strict evaluator. Anmolspace’s public Spark-X2.5-1.7B case study includes an intermediate response that derived 5/33 correctly but placed its ANSWER label inline with the derivation. The scorer required one standalone answer line.
The study contains 24 original development cases: six arithmetic families, two numerical variants, and neutral or misleading-hint prompts for each variant. Its preserved CPU non-thinking batch scored 19/24; a native-thinking GPU batch scored 23/24; a new GPU batch with one uniform answer-format instruction scored 24/24. These are distinct archived batches, not successful answers pooled across runs.
An example makes the reasoning issue concrete. A car travels 120 km at 30 km/h and returns 120 km at 60 km/h. Total distance is 240 km and total time is six hours, so average speed is 40 km/h. The baseline divided only the one-way distance by the total time and answered 20. The final batch used the full round-trip distance.
For this video, we inspected the public scoring script and reran it locally on all three saved-output batches. The results matched the archived score JSON. A separate exact-fraction calculation checked the answer keys and scoring decisions; both GPU manifests also passed their recorded size and hash checks. This is archived-output verification. We did not rerun model inference or independently decode token IDs.
The scope matters: these were development cases revisited during refinement, not a held-out benchmark. Thinking mode, sampling, token budget, batching and runtime changed from the baseline. The score differences cannot be attributed to one isolated change or advertised as general mathematical accuracy.
The post discloses that Codex designed, ran and checked the experiments, with Anmol’s authorization. In the source reviewed for this video, the author’s complete human review was unconfirmed. This showcase presents a community case; it does not determine eligibility or announce an award. The organizer allows disclosed AI assistance and requires human review and accountability.
Try your own reproducible math evaluation with Spark-X2.5-1.7B or 4B. Preserve prompts, raw outputs, model revision, runtime settings and scoring code. Publish the full results in the corresponding model’s Hugging Face Discussions with HER Hack-Astron #6 in the title, then reply to Issue #9 with the direct link. The deadline is September 13, 2026, 24:00 Beijing time (UTC+8). See the event for all requirements.
Community case: https://huggingface.co/XHToken/Spark-X2.5-1.7B/discussions/18
Rules and entry: https://github.com/XHToken/Spark-X2.5/issues/9
AI assistance and human responsibility: https://github.com/XHToken/Spark-X2.5/issues/9#issuecomment-5578566842
Public evidence archive: https://huggingface.co/spaces/Anmolspace/spark-math-eval-20260908/blob/main/spark_math_submission_v4.zip
Open-source project: https://github.com/XHToken/Spark-X2.5
The video uses explanatory graphics recreated from public archived data, not screenshots of a model run performed for this video.
Music: “Impromptu” by Henryk Koman; synthesized piano by Scores2read. IMSLP, CC0 1.0.
https://imslp.org/wiki/Impromptu_(Koman,_Henryk)#IMSLP1051415
https://creativecommons.org/publicdomain/zero/1.0/
Top comments (0)