DEV Community

Bruno Fernandes
Bruno Fernandes

Posted on Fully Autonomous

Cents Matter: Does a Model Notice When One Financial Fact Changes?

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge.

What I Benchmarked

A financial check can fail while looking almost right. A model might recognize an error but give the wrong correction, round at the wrong stage, or confidently decide that two identical amounts are duplicate payments without the evidence to support that conclusion.

I have a degree in Accounting, and I wanted a small, inspectable test of these distinctions. Cents Matter / Centavos Importam contains 40 original synthetic cases written in Brazilian Portuguese. It tests exact monetary reasoning under explicitly stated rules, rather than knowledge of regulations or tax law.

The cases form 20 minimal pairs. Each pair keeps the scenario and recorded amount fixed while changing one material fact. A source locale changes from pt-BR to en-US; a tariff changes from an outflow to a returned fee; two notifications either share a capture ID or represent distinct captures. The correct answer must change with that fact.

The model reports three fields:

{"status":"erro","esperado_centavos":1013,"diferenca_centavos":-1}
Enter fullscreen mode Exit fullscreen mode

That example corresponds to a recorded R$ 10.12 and a correct R$ 10.13: the discrepancy is recorded minus expected, so it is negative. The model must get all three fields right to pass the case. Where evidence is insufficient, the correct response is indeterminado and two nulls, not a guessed amount.

Family What the eight cases test
Numeric locale Decimal and grouping separators under an explicit source convention
Rounding and allocation Line versus total rounding, half-up versus half-even, negative ties and remainder cents
Event identity Capture/refund identifiers, failed attempts and cancelled invoices
Signed cash flows Fee direction, discount order and whether a refund already includes shipping
Evidence sufficiency Missing fees, unaccepted quotes, unidentified captures and missing contractual exchange rates

There are 18 correct records, 18 demonstrably incorrect records and four cases requiring abstention. The model does not see case IDs, pair membership or the answer key. Each case gets a separate conversation.

Scoring is deterministic. Reference arithmetic uses Python Decimal and integer cents. The answer key includes derivations. No second model judges the first model's answer. Structured outputs are parsed by the Kaggle SDK; this benchmark is about the parsed financial decision, not raw JSON typography.

Models Tested

I tested three models through Kaggle Benchmarks SDK 0.6.1: Gemini 3.7 Flash (google/gemini-3.7-flash), gpt-oss-20b (openai/gpt-oss-20b) and Claude Haiku 4.5 (anthropic/claude-haiku-4-5@20251001). This is a deliberately limited comparison across three available model families, including an open-weight model. It does not represent every provider's strongest model.

All three completed the same 40 cases on October 7, 2026, with zero execution errors. Each case received one attempt, no calculator or search tool, requested temperature zero and requested seed zero. Reasoning used provider defaults; support for those controls can differ. This compares the configurations exposed by Kaggle, rather than equalized reasoning budgets.

The table uses the registered task's runs. Before registration, I used a six-case Gemini pilot to check the pipeline and one full interactive Gemini check. Saving the task triggered a full Gemini run again; the registered and interactive runs returned identical parsed answers. No cases, prompts or reference answers were changed after observing model outputs, and no best-of-run selection was made. The other models were added to that frozen task.

Dataset 1.0.0, SHA-256: 0db671bf5687ec711caa7e55b31d0e61ce2fe374fb029070d44c0a6ce8746d3c.

Findings

Model Exact cases Both members of a pair correct Correct required abstention
Gemini 3.7 Flash 39/40 (97.5%) 19/20 (95%) 4/4
gpt-oss-20b 38/40 (95%) 18/20 (90%) 4/4
Claude Haiku 4.5 19/40 (47.5%) 2/20 (10%) 4/4

Case accuracy: Gemini 97.5%, gpt-oss-20b 95%, Claude Haiku 47.5%. Complete-pair accuracy: 95%, 90%, 10%, respectively.

1. A correct warning can still contain the wrong correction. All three missed cashflow-01-B. The scenario has a completed R$159.90 receipt, a R$30.00 refund and a R$2.99 fee movement. In variant A the fee leaves the balance. In B it is returned and enters the balance. The recorded net remains R$126.91.

The correct B calculation is 15990 - 3000 + 299 = 13289 cents, with discrepancy 12691 - 13289 = -598 cents.

Answer for variant B Status Expected cents Discrepancy cents
Reference erro 13289 -598
Gemini erro 12990 -299
gpt-oss-20b ok 12691 0
Claude Haiku ok 12691 0

Gemini recognized that the record was wrong, but its answer matched the receipt minus refund while omitting the returned fee. The other two returned the same answer as variant A. This is why the benchmark checks the proposed amount and discrepancy, not just the warning label. These observations describe the outputs; they do not reveal the models' internal reasoning.

2. Pair scores expose answers that fail to follow the changing fact. Haiku passed 19 individual cases but only two complete pairs. Fifteen pairs had one correct member, three had neither, and eight received identical parsed answers on both sides. gpt-oss-20b repeated an answer across two pairs; Gemini did so across none. Pair accuracy is a stricter complement to case accuracy, not an independent set of 20 extra tests.

The second shared failure was numeric locale. In locale-04-A, the quoted unit price is 9,876, explicitly in en-US notation, with quantity three and a 12.5% discount. The correct total is 9876 × 3 × 0.875 = 25924.50 BRL, or 2,592,450 cents. The recorded total is only 2,592 cents. GPT OSS accepted that record; Haiku returned 2,600 cents. Both passed the paired pt-BR case. Gemini passed both. The scenarios explicitly say that the source convention applies only to the quoted price, so the discount's notation should not change with it.

3. Returning all the fields does not guarantee consistency. Haiku produced three internally inconsistent answers: an ok status with a nonzero discrepancy, or an erro status with zero discrepancy. These were parsed structured responses, not JSON parsing failures. The deterministic validator counted them as failed cases. Haiku also returned some correct expected amounts alongside wrong discrepancies. For example, in evidence-04-B it correctly returned 51,000 expected cents but a discrepancy of -10 instead of -1,000. Field-level correctness can therefore overstate complete-answer usefulness.

4. All models abstained correctly when evidence really was missing. Each passed the four deliberately underspecified cases. However, Haiku also abstained on duplicates-01-A, where the shared capture ID made the answer determinate. Correct abstention on the missing-information subset should not be read as perfect evidence handling, or proof that a model will avoid guessing in less direct real-world documents.

Family (8 cases each) Gemini gpt-oss-20b Claude Haiku
Numeric locale 8/8 7/8 4/8
Rounding and allocation 8/8 8/8 2/8
Event identity 8/8 8/8 4/8
Signed cash flows 7/8 7/8 3/8
Evidence sufficiency 8/8 8/8 6/8

I recomputed the scores from all 120 registered-run responses and checked full coverage, dataset version and prompt hashes. No failed cases were dropped. The pipeline refuses to issue a complete aggregate score for a service or parsing error.

This is a small, synthetic, deliberately constructed test, not a representative sample of accounting work. A 39-versus-38 result is a one-case difference and should not be treated as a robust general ranking. One registered run per model, with the additional Gemini pipeline checks described above, does not estimate variability. Public answer keys can also contaminate future evaluations. High scores here support a narrow conclusion about these explicit cases, not professional competence or production readiness.

A useful next version would add more independent pairs, repeat runs with a fixed protocol and compare tool-free responses with a verified calculator. The shared fee-direction failure suggests testing whether models can first extract signed cash-flow events before performing arithmetic.

My Benchmark

Open the Cents Matter / Centavos Importam benchmark collection.

The collection contains task version 1. Its source includes every prompt, reference derivation and the deterministic scorer. The execution notebooks include predictions and metrics in their Output tabs:

Implementation uses the official Kaggle Benchmarks SDK, including its documented nested-task evaluation pattern. To inspect a claim, find its case ID in cases-*.json or cases-*.csv, compare the three parsed answer fields with the reference, and verify the run's dataset hash in metrics-*.json.

Top comments (1)

Collapse
 
deanlee profile image
Dean Lee •

The gap between marginal case accuracy and complete-pair accuracy isolates why synthetic benchmarks routinely overstate automation readiness in accounting. When Haiku hits 47.5% on raw cases but collapses to 10% on complete pairs, it exposes that the model is often outputting a plausible baseline rather than conditioning on the counterfactual shift. Repeating parsed answers across paired variants means the agent is insensitive to the underlying financial state change.

The asymmetry in error handling is the real operational bottleneck. Flagging an entry as an error while miscalculating the expected amount does not prevent a reconciliation failure; it just introduces a phantom balancing entry into the ledger. In double-entry bookkeeping, an incorrect automated correction carries the exact same audit cost as an undetected error.