This is a submission for the Kaggle Benchmarking Challenge.
What I Benchmarked
A financial check can fail while looking almost right. A model might recognize an error but give the wrong correction, round at the wrong stage, or confidently decide that two identical amounts are duplicate payments without the evidence to support that conclusion.
I have a degree in Accounting, and I wanted a small, inspectable test of these distinctions. Cents Matter / Centavos Importam contains 40 original synthetic cases written in Brazilian Portuguese. It tests exact monetary reasoning under explicitly stated rules, rather than knowledge of regulations or tax law.
The cases form 20 minimal pairs. Each pair keeps the scenario and recorded amount fixed while changing one material fact. A source locale changes from pt-BR to en-US; a tariff changes from an outflow to a returned fee; two notifications either share a capture ID or represent distinct captures. The correct answer must change with that fact.
The model reports three fields:
{"status":"erro","esperado_centavos":1013,"diferenca_centavos":-1}
That example corresponds to a recorded R$ 10.12 and a correct R$ 10.13: the discrepancy is recorded minus expected, so it is negative. The model must get all three fields right to pass the case. Where evidence is insufficient, the correct response is indeterminado and two nulls, not a guessed amount.
| Family | What the eight cases test |
|---|---|
| Numeric locale | Decimal and grouping separators under an explicit source convention |
| Rounding and allocation | Line versus total rounding, half-up versus half-even, negative ties and remainder cents |
| Event identity | Capture/refund identifiers, failed attempts and cancelled invoices |
| Signed cash flows | Fee direction, discount order and whether a refund already includes shipping |
| Evidence sufficiency | Missing fees, unaccepted quotes, unidentified captures and missing contractual exchange rates |
There are 18 correct records, 18 demonstrably incorrect records and four cases requiring abstention. The model does not see case IDs, pair membership or the answer key. Each case gets a separate conversation.
Scoring is deterministic. Reference arithmetic uses Python Decimal and integer cents. The answer key includes derivations. No second model judges the first model's answer. Structured outputs are parsed by the Kaggle SDK; this benchmark is about the parsed financial decision, not raw JSON typography.
Models Tested
I tested three models through Kaggle Benchmarks SDK 0.6.1: Gemini 3.7 Flash (google/gemini-3.7-flash), gpt-oss-20b (openai/gpt-oss-20b) and Claude Haiku 4.5 (anthropic/claude-haiku-4-5@20251001). This is a deliberately limited comparison across three available model families, including an open-weight model. It does not represent every provider's strongest model.
All three completed the same 40 cases on October 7, 2026, with zero execution errors. Each case received one attempt, no calculator or search tool, requested temperature zero and requested seed zero. Reasoning used provider defaults; support for those controls can differ. This compares the configurations exposed by Kaggle, rather than equalized reasoning budgets.
The table uses the registered task's runs. Before registration, I used a six-case Gemini pilot to check the pipeline and one full interactive Gemini check. Saving the task triggered a full Gemini run again; the registered and interactive runs returned identical parsed answers. No cases, prompts or reference answers were changed after observing model outputs, and no best-of-run selection was made. The other models were added to that frozen task.
Dataset 1.0.0, SHA-256: 0db671bf5687ec711caa7e55b31d0e61ce2fe374fb029070d44c0a6ce8746d3c.
Findings
| Model | Exact cases | Both members of a pair correct | Correct required abstention |
|---|---|---|---|
| Gemini 3.7 Flash | 39/40 (97.5%) | 19/20 (95%) | 4/4 |
| gpt-oss-20b | 38/40 (95%) | 18/20 (90%) | 4/4 |
| Claude Haiku 4.5 | 19/40 (47.5%) | 2/20 (10%) | 4/4 |
1. A correct warning can still contain the wrong correction. All three missed cashflow-01-B. The scenario has a completed R$159.90 receipt, a R$30.00 refund and a R$2.99 fee movement. In variant A the fee leaves the balance. In B it is returned and enters the balance. The recorded net remains R$126.91.
The correct B calculation is 15990 - 3000 + 299 = 13289 cents, with discrepancy 12691 - 13289 = -598 cents.
| Answer for variant B | Status | Expected cents | Discrepancy cents |
|---|---|---|---|
| Reference | erro | 13289 | -598 |
| Gemini | erro | 12990 | -299 |
| gpt-oss-20b | ok | 12691 | 0 |
| Claude Haiku | ok | 12691 | 0 |
Gemini recognized that the record was wrong, but its answer matched the receipt minus refund while omitting the returned fee. The other two returned the same answer as variant A. This is why the benchmark checks the proposed amount and discrepancy, not just the warning label. These observations describe the outputs; they do not reveal the models' internal reasoning.
2. Pair scores expose answers that fail to follow the changing fact. Haiku passed 19 individual cases but only two complete pairs. Fifteen pairs had one correct member, three had neither, and eight received identical parsed answers on both sides. gpt-oss-20b repeated an answer across two pairs; Gemini did so across none. Pair accuracy is a stricter complement to case accuracy, not an independent set of 20 extra tests.
The second shared failure was numeric locale. In locale-04-A, the quoted unit price is 9,876, explicitly in en-US notation, with quantity three and a 12.5% discount. The correct total is 9876 × 3 × 0.875 = 25924.50 BRL, or 2,592,450 cents. The recorded total is only 2,592 cents. GPT OSS accepted that record; Haiku returned 2,600 cents. Both passed the paired pt-BR case. Gemini passed both. The scenarios explicitly say that the source convention applies only to the quoted price, so the discount's notation should not change with it.
3. Returning all the fields does not guarantee consistency. Haiku produced three internally inconsistent answers: an ok status with a nonzero discrepancy, or an erro status with zero discrepancy. These were parsed structured responses, not JSON parsing failures. The deterministic validator counted them as failed cases. Haiku also returned some correct expected amounts alongside wrong discrepancies. For example, in evidence-04-B it correctly returned 51,000 expected cents but a discrepancy of -10 instead of -1,000. Field-level correctness can therefore overstate complete-answer usefulness.
4. All models abstained correctly when evidence really was missing. Each passed the four deliberately underspecified cases. However, Haiku also abstained on duplicates-01-A, where the shared capture ID made the answer determinate. Correct abstention on the missing-information subset should not be read as perfect evidence handling, or proof that a model will avoid guessing in less direct real-world documents.
| Family (8 cases each) | Gemini | gpt-oss-20b | Claude Haiku |
|---|---|---|---|
| Numeric locale | 8/8 | 7/8 | 4/8 |
| Rounding and allocation | 8/8 | 8/8 | 2/8 |
| Event identity | 8/8 | 8/8 | 4/8 |
| Signed cash flows | 7/8 | 7/8 | 3/8 |
| Evidence sufficiency | 8/8 | 8/8 | 6/8 |
I recomputed the scores from all 120 registered-run responses and checked full coverage, dataset version and prompt hashes. No failed cases were dropped. The pipeline refuses to issue a complete aggregate score for a service or parsing error.
This is a small, synthetic, deliberately constructed test, not a representative sample of accounting work. A 39-versus-38 result is a one-case difference and should not be treated as a robust general ranking. One registered run per model, with the additional Gemini pipeline checks described above, does not estimate variability. Public answer keys can also contaminate future evaluations. High scores here support a narrow conclusion about these explicit cases, not professional competence or production readiness.
A useful next version would add more independent pairs, repeat runs with a fixed protocol and compare tool-free responses with a verified calculator. The shared fee-direction failure suggests testing whether models can first extract signed cash-flow events before performing arithmetic.
My Benchmark
Open the Cents Matter / Centavos Importam benchmark collection.
The collection contains task version 1. Its source includes every prompt, reference derivation and the deterministic scorer. The execution notebooks include predictions and metrics in their Output tabs:
-
Gemini registered run, execution
20261007T220903366536Z. -
gpt-oss-20b registered run, execution
20261007T221533122434Z. -
Claude Haiku registered run, execution
20261007T221511948876Z.
Implementation uses the official Kaggle Benchmarks SDK, including its documented nested-task evaluation pattern. To inspect a claim, find its case ID in cases-*.json or cases-*.csv, compare the three parsed answer fields with the reference, and verify the run's dataset hash in metrics-*.json.

Top comments (1)
The gap between marginal case accuracy and complete-pair accuracy isolates why synthetic benchmarks routinely overstate automation readiness in accounting. When Haiku hits 47.5% on raw cases but collapses to 10% on complete pairs, it exposes that the model is often outputting a plausible baseline rather than conditioning on the counterfactual shift. Repeating parsed answers across paired variants means the agent is insensitive to the underlying financial state change.
The asymmetry in error handling is the real operational bottleneck. Flagging an entry as an error while miscalculating the expected amount does not prevent a reconciliation failure; it just introduces a phantom balancing entry into the ledger. In double-entry bookkeeping, an incorrect automated correction carries the exact same audit cost as an undetected error.