This is a submission for the Kaggle Benchmarking Challenge.
What I Benchmarked
An assistant can explain sample rates beautifully and still return the wrong number of frames. That gap matters when a number becomes an export setting, a MIDI position, or a file-size estimate.
Audio Arithmetic asks a narrow question: can a model return the exact integer required by an explicitly specified music-workflow calculation?
The benchmark contains 24 original synthetic cases, with four cases in each family:
| Family | Calculation |
|---|---|
| Loop frames | Bars, time signature and quarter-note BPM to sample frames |
| PCM bytes | Frames × channels × stored bits per sample / 8 |
| Resampling | Frame count × target rate / source rate |
| MIDI ticks | Absolute tick × target PPQ / source PPQ |
| Delay frames | Milliseconds × sample rate / 1,000 |
| Trim frames | Total frames minus head and tail trims |
Every prompt defines the relevant units. A frame means one time step across all channels. PCM questions exclude headers and padding. Rounding happens once, at the end, with exact .5 ties rounded upward. This removes several ambiguities that can make an arithmetic evaluation unfair.
The ground-truth generator uses Python's Fraction, so an approximate floating-point intermediate cannot silently change a boundary answer. The complete case set was generated and checked before the full model comparisons. It was not changed after seeing the results.
Each question starts a fresh model chat. The model receives the problem and an integer output schema; it does not receive the expected answer. No external calculator or other tool is supplied.
The score is schema-parsed exact integer accuracy, with one point per correct answer and no partial credit. This is deliberately a calculation test. It does not test listening, audio quality, plugin behavior, or the ability to operate a DAW.
Models Tested
I tested three models available through Kaggle Benchmarks on October 8, 2026, using SDK 0.6.1:
| Exact model identifier | Initial notebook run | Platform task v1 run |
|---|---|---|
google/gemini-3.7-flash |
24/24 (100%) | 24/24 (100%) |
anthropic/claude-haiku-4-5@20251001 |
14/24 (58.3%) | 16/24 (66.7%) |
openai/gpt-oss-20b |
22/24 (91.7%) | 24/24 (100%) |
This lineup compares different model families, including GPT-OSS, on a small task that does not require a specialist dataset. It is not a comparison between equally sized models or equally recent releases.
These are one initial notebook run and one platform task v1 run per model, on the same 24 cases. All six runs completed. The platform leaderboard shows the task v1 results; the initial notebook observations remain separately labeled. All 144 observations were saved with the prompt, expected integer, parsed answer, raw assistant payload and correctness flag. Every saved raw JSON integer matches its parsed value, and the totals were independently recomputed against the identical frozen case set.
Temperature zero was requested, but all three Kaggle model interfaces reported temperature support as false in both contexts. The request therefore does not establish deterministic decoding. Provider defaults apply, and reasoning budgets were not normalized.
Findings
The category breakdown tells more than the total
The following category table and error examples describe the initial notebook runs.
| Category | Gemini 3.7 Flash | Claude Haiku 4.5 | GPT-OSS 20B |
|---|---|---|---|
| Loop frames | 4/4 | 0/4 | 2/4 |
| PCM bytes | 4/4 | 3/4 | 4/4 |
| Resampling | 4/4 | 0/4 | 4/4 |
| MIDI ticks | 4/4 | 4/4 | 4/4 |
| Delay frames | 4/4 | 3/4 | 4/4 |
| Trim frames | 4/4 | 4/4 | 4/4 |
All three models handled the MIDI conversion and trimming cases correctly. Loop duration exposed errors in both Claude and GPT-OSS. That suggests a useful next experiment: expand chained rate-and-duration calculations while retaining simpler subtraction and ratio cases as controls. Four cases per category are too few to establish a general capability hierarchy.
A valid integer can encode the wrong audio format
One question specifies 480,001 frames, six channels and tightly packed 24-bit samples. The payload is:
480001 × 6 × 24 / 8 = 8,640,018 bytes
Claude returned 5,760,012. That number is exactly 480001 × 6 × 16 / 8: the payload that 16-bit samples would occupy. This is a numerical correspondence, not an observed explanation of the model's reasoning. The saved assistant payload exposes only the final JSON answer; it does not provide a reasoning trace.
The output had a valid schema and an integer value. Neither guaranteed that the stated bit depth had been respected.
One frame can be the whole failure
Another question resamples 96,001 frames from 96 kHz to 48 kHz. Under the stated rule:
96001 × 48000 / 96000 = 48000.5 → 48,001 frames
Claude returned 48,000, losing one frame at the exact half-way boundary. Gemini and GPT-OSS returned 48,001.
It would be tempting to label this a consistent rounding preference. The evidence does not support that: Claude correctly rounded another .5 tie in the MIDI category. A single failure identifies a case to investigate, not a universal habit.
A high aggregate score can hide larger timing errors
GPT-OSS answered 22 questions correctly. Its two misses were loop lengths:
- Nine bars of 7/8 at 127 quarter-note BPM, exported at 44.1 kHz: expected
656,291, returned656,309— 18 frames high. - Three bars of 5/4 at 73.5 quarter-note BPM, exported at 96 kHz: expected
1,175,510, returned1,175,612— 102 frames high.
Those discrepancies exceed the rounding boundary itself. Because the saved assistant payloads expose only the final integer, I cannot attribute them to a particular intermediate step.
A second execution changed two scores
The platform task v1 runs kept the same 24 prompts and gold answers. Gemini remained at 24/24; Claude improved from 14/24 to 16/24; GPT-OSS improved from 22/24 to 24/24. Claude still missed all four loop cases and three resampling cases in the platform run.
This is an observation across two execution contexts, not a controlled experiment isolating stochastic variation. It does not establish whether decoding, execution context or another platform detail caused the changes. It does show why publishing only one favorable run would give an incomplete picture. Even a score of 100% in the current leaderboard can coexist with earlier mistakes on the identical questions.
What this changes about using an assistant
The useful distinction is between producing a plausible setting and producing a verified setting. In a workflow that applies these numbers, I would have the assistant extract the parameters, then compute and validate the final value with deterministic arithmetic before applying it. This benchmark provides small, inspectable regression cases for that boundary.
Gemini's two perfect runs are encouraging within this scope. Two runs do not establish long-term repeatability, musical judgment, or reliable operation of an entire audio workflow. It may also mean these 24 explicit questions are already too easy to distinguish stronger models.
The SDK's integer schema can coerce some raw numeric strings or integral floats. The score therefore measures the parsed value, not strict raw-text formatting. In all six actual runs the saved values were raw JSON integers, but the metric itself permits more than that. API and parsing failures abort a run; partial observations are preserved and must not be presented as a complete accuracy score.
Next, I would collect more repetitions, add parameter variations and larger boundary sets, and compare direct answers with a calculator-enabled condition. Those are proposed extensions, not results from this entry. I would keep the current cases fixed as a published baseline so that later changes remain distinguishable.
My Benchmark
Audio Arithmetic: 24 Music Workflow Checks on Kaggle
The backing task notebook contains the original prompts, rational ground-truth generator, scoring function and incremental observation export. The description distinguishes the initial notebook observations from the platform task v1 evaluations, and the leaderboard displays the latter.
This project was created with AI agents, including implementation, validation, analysis and article drafting. The task uses no private recordings, user files or third-party dataset. The execution framework and model access come from Kaggle Benchmarks and its official Python SDK.
Top comments (0)