Change the numbers while keeping the problem structure. Does the model calculate a new answer? In jxyytf’s community case, Spark-X2.5-1.7B was tested on six original math problems, each paired with a numeric variant. The participant reports 12/12 correct, with one sample per problem and exact numeric scoring.
One pair shows the idea
The example in the post changes a bead-counting problem:
| Version | Calculation | Expected answer |
|---|---|---|
| Original | 14 + 9 − 6 | 17 |
| Numeric variant | 18 + 11 − 7 | 22 |
The structure stays fixed while the quantities change. This gives a concrete way to probe whether a model follows the new numbers. A few matched items cannot establish broad robustness or rule out memorization.
What the participant reports
The run used Spark-X2.5-1.7B Q4_K_M through the XHToken llama.cpp fork on a macOS arm64 CPU, with four threads and zero GPU layers. The author reports no paid cloud resources. The six original problems comprise four arithmetic/word problems and two algebra/geometry problems, each with one numeric variant. They are newly authored items, rather than official GSM8K or MATH benchmark splits.
The reported results are 6/6 originals and 6/6 variants. Each prompt received one sample: pass@1, without majority voting. The scorer uses exact numeric matching after extracting an answer, counts unparseable outputs as incorrect, and preserves both reasoning_content and content.
These are participant-reported results. We did not rerun the model. The post says the raw records and scorer were retained, but as of September 9, 2026 it does not provide public download links to the complete artifacts. No failures were reported in this 12-item slice. The small sample and related item pairs limit the conclusions. This showcase does not establish official acceptance or announce a winner.
Bring your own math experiment
HER Hack-Astron #6 is open. Test Spark-X2.5 1.7B or 4B with a question you can investigate: change numbers, reword a problem, inspect a wrong reasoning step, or compare a Python-assisted run with a no-tool baseline.
- Pin the model revision and document your dataset, prompts, hardware, runtime and decoding settings.
- Retain and publish the raw outputs, scoring code and representative reasoning, including failures. Distinguish final-answer accuracy from reasoning quality.
- Publish the complete evaluation in the tested model’s Hugging Face Discussions. Include HER Hack-Astron #6 in the title.
- Reply to event Issue #9 with the direct Discussion link.
Deadline: September 13, 2026, 24:00 Beijing time (UTC+8), meaning September 14 at 00:00. Award: one winner receives USD 100. For team entries, at least 50% of listed contributors must have profiles identifying them as women. Read the full rules for all acceptance requirements.
Watch the short · Read the case · Open-source project
Video music: Azul, composed and performed by Akiana Molina, from IMSLP, licensed under CC BY 4.0. Excerpted, level-adjusted and mixed under narration.
Top comments (0)