This is a submission for the Kaggle Benchmarking Challenge
An AI model can classify hundreds of support tickets, organize them into categories, and produce a summary. But do the numbers in that summary actually agree with the records it just listed?
That question led me to build Cross-View Conservation, a benchmark that tests whether a model can keep two representations of the same answer consistent as the workload grows.
The most interesting finding was that getting individual items right did not guarantee that the final summary added up.
What I Benchmarked
Cross-View Conservation evaluates deterministically generated synthetic support tickets at three workload sizes: 100, 400, and 800 tickets. I use three fixed data seeds at each size, giving nine cases per model.
Each ticket contains fields such as its region, customer tier, service, severity, and status. The model receives explicit classification rules, applied in this order:
- If
status=closed, assign the ticket toCLOSED. - Otherwise, assign it to
ESCALATEif the customer tier is enterprise and severity is at least 3, or if the service is security and severity is at least 4. - Otherwise, assign it to
NORMALif the customer tier is business and severity is at least 2, or the region is EU and severity is at least 2. - All remaining tickets go to
QUEUE.
For example, this synthetic ticket belongs to CLOSED because the first rule takes priority:
TCK-0001 | region=APAC | tier=enterprise | service=search | severity=5 | status=closed
Each model must return two views of the same classification:
-
View A: Detailed partitions. List ticket IDs under
ESCALATE,NORMAL,QUEUE, andCLOSED. - View B: Aggregate summary. Report each category's ticket count and total severity, together with the overall ticket count.
The model does not receive the hidden reference labels. It must classify the records and generate its own summary.
The severity totals require adding the severity values of tickets in each category. At 800 tickets, that means potentially hundreds of additions. The prompt provides no calculator, code execution, or external tools. The benchmark therefore tests rule application, ID tracking, long addition, and consistency between two representations of an answer.
My primary metric is exact self-consistency. The evaluator recomputes counts and severity sums from the model's own lists and compares them with the reported summary. A mismatch of even one unit in a required total fails the case.
Self-consistency is not the same as correctness. A summary can agree with an incorrectly classified list. That is why I also track item coverage, duplicate assignments, classification accuracy against a deterministic reference, count agreement, and severity-total agreement.
I wanted to separate these failure modes instead of hiding them inside one accuracy score.
Models Tested
I chose six models across Anthropic, OpenAI, and Google to compare behavior across providers and different model offerings. The lineup includes two Gemini Flash versions, making their results particularly interesting to compare under the same task design.
The current Kaggle leaderboard shows these results:
| Model | Cases passed | Score |
|---|---|---|
| Claude Opus 5 | 9/9 | 100.0% |
| Claude Haiku 5.5 | 7/9 | 77.8% |
| GPT-6 Luna | 5/9 | 55.6% |
| Gemini 3.8 Flash | 4/9 | 44.4% |
| Gemini 3.7 Flash | 3/9 | 33.3% |
| GPT-5.4 mini | 0/9 | 0.0% |
Each case contributes one ninth of the score. The task code averages all nine cases and assigns zero to caught execution failures. Under that scoring logic, Claude Opus 5's displayed score of 1.00 corresponds to 9/9.
There is an important distinction between official results and my earlier validation. In that separate validation run, five Claude Opus 5 cases completed and passed, while four calls were blocked by provider quota. Those validation results are not the official Kaggle score.
Evaluation settings
Every case ran in a fresh conversation using llm.prompt(prompt). I did not supply external tools, set an explicit reasoning override, specify a custom output-token cap, or set a model-generation seed. Provider and model defaults applied. The data seeds were fixed.
This is an end-to-end comparison under the available default settings, not a controlled comparison of reasoning effort or token efficiency. Models can spend different numbers of output tokens.
The current leaderboard displays GPT-6 Luna, while an earlier validation log identifies openai/gpt-5.6-luna. I preserve the logged identifier when discussing that validation finding rather than assume the two labels are interchangeable.
Findings
1. Correct classifications do not guarantee correct aggregation
In one 100-ticket validation case, the model logged as openai/gpt-5.6-luna classified all 100 tickets correctly, yet its reported QUEUE severity total was off by 24.
The individual classifications were correct in that case, but the aggregate summary still contradicted the records.
A classification-only benchmark could mark the labels correct and miss this inconsistency entirely.
2. Matching counts can still hide incorrect severity totals
Gemini 3.7 Flash produced another revealing pattern in validation.
At 400 tickets, it classified all 400 tickets correctly in all three seeds, and the category counts matched. Yet its severity totals were wrong in every case, with absolute discrepancies ranging from 1 to 19.
These are not necessarily large numerical errors. But exact-match evaluation is intentionally strict: even a small discrepancy means the summary no longer agrees precisely with the detailed result.
At 800 tickets, category counts drifted as well. In the observed validation cases, the QUEUE count was off by between 4 and 8.
The saved October 10 official Kaggle run is consistent with this pattern: all three 800-ticket Gemini 3.7 Flash cases failed the exact aggregate-consistency checks.
Three cases at each workload size are not enough to establish a universal scaling law. More repeated evaluations are needed to determine how stable the pattern is.
3. Duplicate assignments can invalidate the partition itself
In one 800-ticket validation run, GPT-5.4 mini listed 710 of the 800 ticket IDs more than once, producing 1,857 ID listings for 800 tickets. The most frequently repeated ID appeared four times.
Its output was not a partition: a valid partition requires each input ticket to appear exactly once across the four categories.
This is different from miscalculating a summary over otherwise valid lists. It also shows why duplicate and coverage diagnostics matter alongside the headline score.
4. A refusal is different from a wrong answer
In two larger-workload validation cases, the model logged as openai/gpt-5.6-luna declined to produce the requested classification and summary. One response stated:
I'm unable to reliably produce the exact classification and summaries for all 800 tickets without risking omissions or misclassification.
A refusal is different from confidently returning incorrect totals. Both affect whether the task was completed, but they reveal different behaviors. That is why I examine individual outputs and execution coverage rather than treating every zero score as the same kind of failure.
5. A related benchmark asks a different question
The closest related work I found is Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens.
That benchmark tests whether models count rows returned by a tool, comparing approaches that return matching IDs with approaches that provide the count directly. Its reported findings discuss a relationship between output-token use and counting accuracy in that particular setup.
Cross-View Conservation asks a different question. It tests whether a model's own ticket partition and its summary of that partition agree as the batch grows. The model generates both views instead of receiving a list from an external tool and being asked to count it.
The two benchmarks are complementary. One studies counting tool-returned rows; the other studies consistency between two representations of a model-generated answer. My results do not establish what happened inside a model's hidden reasoning, nor do they show that spending more tokens necessarily fixes the problem.
What I Would Measure Next
This is an initial benchmark, not a definitive ranking of model quality. Nine cases per model, variable outputs, provider quotas, and differences in provider defaults all limit what can be concluded.
I would prioritize three follow-up experiments:
- Increase repetitions. Use more seeds and report completed, failed, and blocked cases separately, with an explicit denominator for every model.
- Vary workload difficulty independently of size. Change category balance, ticket ambiguity, and severity distributions separately to identify conditions associated with failures.
- Compare generation strategies. Test whether asking a model to produce the detailed lists first and calculate the summary in a separate step improves consistency compared with producing both views together.
I would also report output-token usage alongside the failure diagnostics. That could help characterize the relationship between output length, consistency, cost, and workload size without assuming that more tokens necessarily mean better reasoning.
The useful next step is not simply to add more models. It is to understand which failure modes recur, under which conditions, and whether a different generation strategy reduces them.
My Benchmark
Explore the task, inspect the results, and compare the models on Kaggle:
Cross-View Conservation: Do AI Models' Answers Still Add Up at Scale?
The idea is simple: an answer should not merely look plausible. When it contains detailed records and a summary of those records, the two views should agree.
Cross-View Conservation turns that expectation into a measurable test and investigates what happens as the workload grows.

Top comments (2)
With nine cases per model the ranking is looser than the table suggests. Wilson 95% intervals: Opus 9/9 is 70-100%, Haiku 7/9 is 45-94%, Luna 5/9 is 27-81%, Gemini 3.8 Flash 4/9 is 19-73%, 3.7 Flash 3/9 is 12-65%, and 0/9 is 0-30%. Fisher exact for Opus vs Luna is p about 0.08, Opus vs Haiku about 0.47, Luna vs Gemini 3.8 about 0.64. Only the two ends (Opus, GPT-5.4 mini) are separated from the pack; the middle four are not ordered by this data.
The nine cases are also not nine independent draws: three seeds at each of three sizes, and a one-unit error fails the case, so pass probability falls with the number of additions. I would report pass rate by size (100/400/800), and ideally the per-category absolute error in severity totals, which would show whether a model is failing by a little at 800 or by a lot. A model that is off by 1 at 800 tickets and one that miscounts at 100 look identical in pass/fail.
One cheap addition since the evaluator already recomputes from the model's own lists: report "consistent but wrong" and "correct but inconsistent" as separate columns. Did any model's failures concentrate in the QUEUE bucket, the one defined by exclusion?
Official Platform Update
Security protocols have been updated for all developer accounts.