This is a submission for the Kaggle Benchmarking Challenge.
A gene can increase in one cell type and decrease in another. Ask whether its average expression increased, and the answer depends on which mixture of cells you mean.
There is a subtler question: can we know that expression increased without knowing its exact percentage increase?
Yes. And a model should be able to say both things at once.
I built CellCaution to test that boundary. In a small, controlled study, Gemma answered every primary case correctly. Verification instructions gave GPT-5.4 mini more correct verdicts but fewer correct complete answers. Claude's strict score looked disastrous until I separated response-format compliance from the mathematical content.
Those findings changed how I would evaluate an AI assistant for scientific analysis: the verdict, the number, and the output contract each need their own evidence.
What I Benchmarked
My interest comes from working on AI for single-cell RNA sequencing. Before asking a model to interpret a complicated biological study, I wanted to know whether it could handle a small descriptive comparison whose answer I could calculate exactly.
Consider these synthetic expression measurements:
| Cell type | Control mean | Treated mean |
|---|---|---|
| A | 17 | 22 |
| B | 53 | 50 |
Type A increased by 5 units. Type B decreased by 3. We standardize both groups to the same reference composition, so the comparison uses the same weights on both sides.
At 80% A and 20% B, the control mean is 24.2 and the treated mean is 27.6. The increase is approximately 14.05%.
Now change only what we know about the reference population:
| Information supplied | Is a strict increase supported? | Unique percentage | Possible percentage range |
|---|---|---|---|
| Exactly 80% A | Yes | 14.05% | 14.05% to 14.05% |
| A is between 65% and 85% | Yes | null |
7.43% to 16.96% |
| Any composition is allowed | Insufficient information | null |
−5.66% to 29.41% |
In the middle row, every allowed reference gives an increase. But the constraints do not identify one percentage. Returning null is the mathematically correct answer to the magnitude question, even though the direction is known.
Figure 1. These are final v2 measurements, not an example drawn from the earlier pilot.
That is the specific behavior I wanted to measure: preserving what is known while refusing to manufacture what is not.
A benchmark with exact answer keys
The primary dataset contains 24 cases: eight templates, each presented under fixed, constrained, and unrestricted reference compositions. The templates cover two or three cell types and four expression patterns: mixed changes, additive increases, proportional changes, and no change.
Each response must provide four JSON fields:
{
"verdict": "supported",
"standardized_change_percent": null,
"percentage_range": [7.4324324324, 16.9642857143],
"explanation": "Every allowed reference gives an increase, but the percentage varies."
}
The verdict is supported when the treated mean is strictly higher at every allowed reference, unsupported when it is higher at none, and insufficient_information when the answer depends on the reference.
The primary score requires both the correct verdict and the correct percentage or null. Range accuracy is measured separately. Numerical answers use a tolerance of 0.05 percentage points. The explanation is retained for inspection, but an LLM judge does not score its prose.
The answer keys use rational arithmetic. Allowed populations are every convex mixture of the supplied reference vertices. Differences are linear in the weights. With positive control means, the percentage ratio at any mixture is a weighted average of the vertex ratios, with weights proportional to each vertex's control mean. Its extrema therefore occur at vertices. This gives exact bounds without relying on a sampled grid or another model's judgment.
Models Tested
I used three models available through the Kaggle Benchmarks registry:
| Model | Requested registry ID |
|---|---|
| Gemma 4 31B | google/gemma-4-31b |
| GPT-5.4 mini | openai/gpt-5.4-mini-2026-03-17 |
| Claude Sonnet 4.6 | anthropic/claude-sonnet-4-6@default |
This was a manageable lineup spanning three model families, including Gemma and two hosted assistants. It is a comparison of these requested models on this task, rather than a claim about their providers as a whole.
For each model I recorded 24 direct responses, 24 responses with explicit verification instructions, and eight reordered controls: 168 included case evaluations. The verification prompt asks the model to calculate weighted means at each vertex and check signs, extrema, and uniqueness before returning its answer.
I used SDK defaults rather than a shared, explicitly controlled temperature or reasoning budget. The exported IDs identify the requested models; they do not independently verify the exact backend revision served.
Findings
1. A score must say what it counts
Here are the direct-prompt results under the frozen output contract. Invalid outputs count as incorrect in these totals.
| Model | Valid output / 24 | Verdict correct / 24 | Percentage or null correct / 24 | Range correct / 24 | Primary correct / 24 |
|---|---|---|---|---|---|
| Gemma 4 31B | 24 | 24 | 24 | 24 | 24 |
| GPT-5.4 mini | 24 | 18 | 16 | 14 | 12 |
| Claude Sonnet 4.6 | 1 | 1 | 1 | 1 | 1 |
Gemma was consistently correct on this small dataset. That is a useful result, but 24 related synthetic cases do not establish broad scientific reliability.
Claude's result needs a different explanation. In 23 of its 24 direct responses, the frozen parser rejected the response format. Extra text around a JSON block violates the requested contract. A whole JSON object or a single enclosing fence is accepted; surrounding commentary is not.
I kept that strict score and added a post-hoc diagnostic, applied uniformly across models: for an invalid response containing exactly one fenced JSON block, extract that block and run the unchanged grader. No values are repaired, and no additional model calls are made.
Under this diagnostic, Claude's direct primary result becomes 23/24. Its remaining error is mathematical: in the fixed-composition example, it returned about 22.40% instead of 14.05%, consistent with averaging cell-type percentage changes rather than comparing the weighted means.
Figure 2. The extraction diagnostic does not replace the frozen strict score.
The practical lesson is to report both layers. A downstream program needs a usable response. A scientific evaluation also needs to distinguish an unusable response from incorrect mathematical content.
2. Verification helped one component and hurt another
The verification instructions produced this primary-score comparison:
| Model | Direct / 24 | Verification / 24 |
|---|---|---|
| Gemma 4 31B | 24 | 24 |
| GPT-5.4 mini | 12 | 9 |
| Claude Sonnet 4.6 | 1 | 0 |
For GPT, verdict accuracy increased from 18 to 21, while percentage-or-null accuracy fell from 16 to 10. The combined primary score fell from 12 to 9. Range accuracy remained 14/24.
Figure 3. One recorded generation per case and prompt arm; this comparison needs replication.
The instruction to verify an answer is not itself evidence of improvement. Here it helped the qualitative decision while the numerical contract became less reliable. A verdict-only benchmark would have missed that tradeoff.
Claude's verification responses were all invalid under the strict contract; all 24 passed the separately reported extraction diagnostic. Again, the strict score alone cannot explain the failure mechanism.
3. Reordering the evidence exposed another instability
For eight constrained cases, I reversed the cell-type rows and every corresponding weight. The mathematical problem and gold answer stay the same.
Gemma preserved its verdict, percentage/null, and range across all 8/8 pairs. GPT preserved all three across 3/8 pairs. Claude also matched across 8/8 under the extraction diagnostic; its reordered strict outputs were invalid, so strict agreement is not assessable.
This is evidence of observed answer instability under equivalent presentations. With one generation per presentation, I cannot separate an ordering effect from ordinary generation variability. The next experiment should repeat both orders with controlled settings.
What I had to fix in the benchmark itself
The first control exports failed the audit: they contained primary-case IDs and verification prompts rather than the intended reordered controls. The execution path had reused results across task contexts.
I excluded those exports and reran the controls under a separate task identity. The included recovery files contain the intended eight control IDs and exact prompts for each model. I then checked all 168 included records against the fixtures and recomputed their strict grades from the raw responses.
That repair matters. An evaluator that checks only whether a run finished can quietly score the wrong experiment.
What changed my thinking
I started by testing whether models could get a biological comparison right. I ended up caring just as much about whether they could represent the boundary of the available evidence.
For an analysis assistant, I would now check three things independently: whether its conclusion follows from the constraints, whether its number is actually identifiable, and whether software can consume its response. A confident explanation cannot substitute for any of those checks.
This study uses synthetic descriptive measurements. It does not test raw sequencing analysis, biological replication, causal treatment effects, significance testing, or clinical decisions. The 24 cases share eight templates, and the verification comparison is exploratory. I would next add repeated runs, more unseen template families, and a calculator-assisted arm to distinguish arithmetic errors from failures to reason about missing information.
My Benchmark
The public leaderboard currently contains Gemma's publication-validation run. The tables in this article describe the separately audited three-model study; they should not be read as three imported leaderboard runs.
The implementation builds on the Kaggle Benchmarks Python library and cookbook. The synthetic fixtures, reference-constraint design, rational answer keys, and component-level evaluation were developed for CellCaution. AI assistance was used for implementation, auditing, visualizations, and editing; I reviewed the measurements and claims against the exported responses.
The question I want a scientific assistant to preserve is simple: what can these measurements establish—and what would it be inventing?
Inspect the audited study evidence
The dataset contains exact prompts, raw model responses, answer keys, recorded scores, and the frozen protocol for all 168 included evaluations. It includes the recovered reordered controls and excludes the invalid original control exports.
One failure, two different problems
In direct case F01, Claude Sonnet 4.6 saw control means of 17 and 53, treated means of 22 and 50, and a fixed reference of 80% type A and 20% type B.
Its returned JSON reported:
{
"verdict": "supported",
"standardized_change_percent": 22.397,
"percentage_range": [22.397, 22.397]
}
Excerpt: the explanation field and surrounding prose are omitted here; the evidence dataset retains the full response.
The verdict was correct. The number was not. The control weighted mean is 24.2 and the treated weighted mean is 27.6, so the requested increase is:
100 × (27.6 − 24.2) / 24.2 = 14.0496%.
The response instead averaged the cell-type percentage changes:
0.8 × [100 × (22 − 17) / 17] + 0.2 × [100 × (50 − 53) / 53] ≈ 22.3973%.
These formulas answer different questions because the cell types have different control baselines. Averaging their percentage changes is not the percentage change of their weighted mean.
This response also included prose outside its JSON block, so it failed the frozen output contract. Extracting the block resolves that formatting issue for the separate diagnostic, but leaves the mathematical error intact. A correct verdict and a detailed explanation did not guarantee a correct magnitude.
This notebook verifies the frozen protocol and exact prompts, recomputes all 168 included evaluations from raw responses, and reproduces the score tables and charts without additional model calls.



Top comments (0)