This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
I tested whether language models can update Excel-style cell references when a formula is copied to a new cell.
The 20 cases covered relative, absolute, and mixed references; ranges; formulas with multiple references; and copying in different directions. Each model was asked to return only the updated formula. I scored an answer as correct when its first line matched the expected formula, ignoring whitespace and letter case.
I started with five simple cases, and all three models scored 1.00. That ceiling result pushed me to add longer copy moves and more complex formulas.
Models Tested
I tested Claude Sonnet 5.5, Claude Opus 5.5, and Gemini 3.7 Flash—the models available to me on Kaggle. The two Claude models let me compare results within one model family; Gemini gave me a comparison across providers.
Findings
Model
Score
Claude Sonnet 5.5
0.25
Claude Opus 5.5
1.00
Gemini 3.7 Flash
1.00
The biggest surprise was what the low score meant. Sonnet often explained how the references should move, but did not put a complete, copy-ready formula on the first line. One case returned the original formula unchanged. Because this task requires a usable formula, those responses failed—even when the explanation suggested the model understood part of the transformation.
That changed how I think about spreadsheet assistance: understanding a formula’s movement is only part of the job. The output also has to be usable in the cell.
Next, I’d separate formula correctness from output-format compliance, then test more Excel edge cases and repeat runs to see whether the model differences hold.
My Benchmark
kaggle.com/benchmarks/anvipardhi/fill-accuracy
Top comments (1)
Does the scorer normalize whitespace and case only outside quoted strings? I'd add a copy fixture such as =A1&"ID A1" and check that the cell reference moves but the literal text stays byte-for-byte unchanged. A second case with two spaces inside a literal would test whether normalization hides an output change.
Separating reference transformation, literal preservation, and first-line format compliance would make the low score easier to interpret. I haven't run the benchmark; these are scorer and copy-boundary fixtures suggested by your scoring rule.