This is a submission for the Kaggle Benchmarking Challenge.
What I Benchmarked
Resolve measures a small code change. I wanted a score for the failure I keep seeing: the edit does the thing that was asked, and drops a behavior the prompt never said again.
Each item is one Python function and one requested edit. The model returns only the new function. The score is how many of 16 items are fully resolved, as a float from 0 to 1. A half-fix scores zero on that item.
There are two tasks. resolve checks the new behavior and the old behavior the prompt never restated. resolve-plus adds an edge suite the prompt never showed. The two tasks are separate generations, so the second score can land above the first.
The result this post is about: GPT-6.1 Sol and GPT-5.6 Luna are the only models at 0.9375 on both tasks. The task page rounds that to 0.94. Several other models match 0.9375 on resolve and then score lower on resolve-plus.
How it works
The prompt shows the function, the one change, and two sample inputs. Anything the model writes besides the new function is thrown away. Hidden tests then run in two lists.
- Fail-to-pass is the new behavior.
- Pass-to-pass is behavior the prompt never repeated.
An item counts only when every test in both lists passes. A timeout, a crash, a missing function, or an empty suite is a failure. One failed test fails the item.
resolve-plus uses the same rule and adds the edge suite. It is a new sample of the 16 items, not a second grade of the first sample.
Inspiration
This is the rule SWE-bench uses for a resolved issue, scaled down to functions I wrote. The paper is Jimenez et al., SWE-bench: Can Language Models Resolve Real-World GitHub Issues? (ICLR 2024).
SWE-bench does not give credit because a patch compiles or because the new test passes. The issue is resolved only when the fail-to-pass tests and the pass-to-pass tests all pass. Resolve keeps that bar and drops the rest of the scaffolding: no scraped repository, no gold patch in the prompt, no tests written by the model. The failure worth reading is an edit that does the asked change and drops an invariant the prompt did not say again.
Models Tested
I started with one model from each of four families, then ran the Claude, GPT, and Gemini lines the proxy would serve, plus the open models listed below. Frontier and open models are charted separately so a single average cannot hide which group is carrying the score.
Twenty-three models have both scores. A model is on a chart only when both have finished. Blue is resolve. Orange is resolve-plus. Charts are sorted by resolve-plus.
Frontier: Claude Haiku 4.5, Haiku 5.5, Sonnet 5, Sonnet 5.5, Opus 5.5, GPT-5.4 mini, GPT-5.6 Luna, GPT-5.6 Sol, GPT-6.1 Sol, Gemini 3.5 Flash, Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemini 3.7 Flash, and Gemini 3.8 Flash.
Open: Gemma 4 31B, Gemma 4 26B, gpt-oss 20B, gpt-oss 120B, DeepSeek-R1, GLM-5, Qwen 3 235B, Qwen 3 Coder 480B, and Qwen 3 Next 80B Instruct.
Grok 4.5 and Grok 4.6 are in the Kaggle model list. The proxy returned model not found for both, so they are not on the charts. Qwen 3 Next 80B Thinking timed out three times on that same proxy, the last time after 3 of 16 items, so it is off the charts too.
Findings
All 23 models
GPT-6.1 Sol and GPT-5.6 Luna are tied at the top, 0.9375 on both tasks. That is 15 of 16 items, twice.
The next group matches the 0.9375 resolve score and lands at 0.875 on resolve-plus: DeepSeek-R1, Claude Sonnet 5.5, and Claude Opus 5.5. Each of them resolved 15 items on the first task and 14 on the second.
gpt-oss 20B is the bar that runs the other way, 0.75 then 0.875. The second task is a new sample, so a higher resolve-plus score is a different generation, not the edge suite helping. The floor of the combined chart is Qwen 3 235B and Gemma 4 26B, both at 0.6875 on resolve-plus.
Frontier models
The frontier chart is the same tie at the top, then a clean step down.
Opus 5.5 and Sonnet 5.5 both open at 0.9375 and close at 0.875. GPT-5.6 Sol opens at the same 0.9375 and closes at 0.8125, a two-item drop. Sonnet 5 holds 0.875 on resolve and falls to 0.75 on resolve-plus.
Haiku 4.5 and Haiku 5.5 sit on the same pair, 0.875 and 0.875. On these 16 functions the newer small model did not move the score. Gemini 3.5 Flash and Gemini 3.6 Flash sit with them. Gemini 3.5 Flash-Lite is the frontier floor, 0.75 on both.
Open models
DeepSeek-R1 leads the open chart at 0.9375 and 0.875, level with Opus 5.5 and Sonnet 5.5 on the frontier chart. Gemma 4 31B is the only open model that holds 0.875 on both tasks. Gemma 4 26B is last, 0.75 then 0.6875.
Qwen 3 Coder 480B, at 0.875 and 0.8125, beats the larger general Qwen 3 235B, at 0.8125 and 0.6875. gpt-oss 120B and Qwen 3 Next 80B Instruct are flat at 0.8125. gpt-oss 20B is the open-chart exception: its second sample scores 0.875 after a 0.75 first sample.
Score against list price
The runs did not store tokens or a dollar bill, so this is not measured cost per item. The horizontal axis is the published output price per million tokens, taken from a rate card last updated 8 October 2026. That card reports the rates against the Anthropic, OpenAI, and Gemini pricing pages. Expensive is on the left. Cheap is on the right. A line connects models from the same provider. Open models are off this chart, and so are Gemini 3.6 Flash and Gemini 3.7 Flash, because that rate card does not list them. Haiku 5.5 is the under-100k rate. These prompts are one short function.
GPT-5.6 Luna is the efficient point: 0.9375 at $1.20 per million output tokens. GPT-6.1 Sol matches that score at $10. Claude Haiku 5.5 holds 0.875 at $0.50, the same score as Haiku 4.5 at $5 and Opus 5.5 at $20. Sonnet 5 and Sonnet 5.5 share the $10 output price. Sonnet 5.5 is the one at 0.875.
What a flat leaderboard hides
The first pass was one model from each of four families: Claude Haiku 4.5, Gemma 4 31B, GPT-5.4 mini, and Gemini 3.7 Flash. Each resolved 14 of 16 items. The leaderboard looks flat until the second row. Haiku and Gemma stay at 0.875 on resolve-plus. GPT-5.4 mini and Gemini 3.7 Flash fall to 0.8125.
Those two drops are different events.
On ring_push, each of those four models, on both tasks, made the requested change and failed a behavior the prompt never restated. Fail-to-pass passed. Pass-to-pass did not. A looser score would have counted those edits. This one does not. That miss was stable across those eight runs.
GPT-5.4 mini’s lower resolve-plus score is the extra suite. On truncate, the requested change and the old behavior both passed, and the edge suite failed. Haiku and Gemma did not lose an item that way.
Gemini 3.7 Flash’s lower resolve-plus score is a different generation of the same items. collapse_name passed the first time and failed the second. The extra suite was not what moved it.
One number is hiding three things: an edit that breaks an unstated invariant, an edit that misses an edge the prompt never showed, and a second sample of the same item that comes out differently.
One generation was enough to move collapse_name from a pass to a fail. A one-shot score is a noisy picture of a model that is close to the line. The next run I would make is more than one sample per item.
My Benchmark
The public benchmark is Resolve on Kaggle. It is version 1, Apache 2.0, and marked public. The leaderboard on that page shows the first four models. The task pages list every run, including ones that did not finish. A model is in the charts here only when both scores finished.
The two tasks:
Suggested citation: Aditya Kumar Puri. Resolve. https://www.kaggle.com/benchmarks/puriadityakumar/resolve/versions/1, 2026.
References
- Jimenez, Carlos E., John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. “SWE-bench: Can Language Models Resolve Real-World GitHub Issues?” ICLR 2024. https://arxiv.org/abs/2310.06770
- SWE-bench. https://www.swebench.com/
- Kaggle Community Benchmarks. https://www.kaggle.com/benchmarks
- kaggle-benchmarks Python library. https://github.com/Kaggle/kaggle-benchmarks
- Google AI. “Introducing Community Benchmarks on Kaggle.” https://dev.to/googleai/introducing-community-benchmarks-on-kaggle-35nc
- Frontier model API pricing, last updated 8 October 2026. https://www.developersdigest.tech/blog/frontier-model-api-pricing-june-2026
- Anthropic pricing. https://www.anthropic.com/pricing
- OpenAI API pricing. https://openai.com/api/pricing
- Gemini API pricing. https://ai.google.dev/gemini-api/docs/pricing






Top comments (0)