This is a submission for the Kaggle Benchmarking Challenge
Those who read my articles know that a lot of them are actually benchmarks of something. Mostly .NET related, but still, someone could say that this challenge should be pretty close to what I usually do.
The opposite was true.
Building a code benchmark and benchmarking AI models are two different worlds, so when I first saw this challenge, I had no idea what exactly I should benchmark. Comparing models on coding tasks felt too generic, and I didn't want to create a benchmark just for the sake of having one.
Then I looked at the topics of my last couple of articles. A lot of them were about APIs, resilience, failures, and how systems behave when something goes wrong. And that gave me an experiment idea: What if I benchmark AI models on one very simple question: To Retry or Not to Retry?
A 503 Service Unavailable does not automatically mean that retrying is safe. A POST request may already have been processed. An idempotency key can completely change the answer. A timeout may happen before the server receives anything or after it has already changed some state.
So the HTTP status code alone is often not enough. The model has to understand the whole situation.
What I Benchmarked
The idea is simple. I prepared several API failure scenarios containing information about the request, the response, and some additional context. The model has to return two things:
- a decision:
YES,NO, orYES_AFTER_DELAY - a one-sentence explanation of why
For example:
POST /payments
503 Service Unavailable
An idempotency key was supplied and the API guarantees
duplicate requests with the same key are not processed twice.
Retry?
A developer will probably immediately say:
YES_AFTER_DELAY
The important part is not only the 503. The context tells us that an idempotency key was supplied and repeated requests with the same key will not process the payment twice. But will an AI model notice the same thing? And more importantly, what happens when the scenario is less obvious?
For this small experiment, I prepared scenarios where the answer depends on details such as HTTP method, status code, idempotency, rate limiting etc.
Every model gets the same response format:
Decision: YES | NO | YES_AFTER_DELAY
Reason: <one sentence>
For scoring, I only use the Decision. This keeps the benchmark simple and deterministic. The Reason does not affect the score. I collect it because it can show whether the model actually understood the scenario or simply arrived at the correct answer for the wrong reason. That also makes the benchmark easy to compare between models. Each scenario has an expected decision, so the final result can simply be calculated as the percentage of retry decisions the model got right.
But why benchmark this at all?
Imagine that you are integrating an external service and want to make the call more resilient. In today's world of AI-assisted coding, there is a very good chance that a coding agent will do the work. The AI can easily generate a retry policy. The more interesting question is: Will it retry the right requests? Because retrying something that should not be retried can be much worse than not retrying at all.
That's what I wanted to measure. Now let's have a look at the scenarios.
Scenarios
I prepared 14 scenarios, each based on an actual task used in the Kaggle benchmark. Originally, all of them were part of the article, but that made it too long for a small experiment. So I decided to move the full scenarios into a standalone app, where you can browse every task together with its expected answer and explanation. And if you want to try them yourself first, there is also a test for humans. After all, why not benchmark some humans too? 😁
Try the human benchmark or browse all scenarios: To Retry or Not to Retry?
Here are the scenarios used in the experiment:
-
Safe Payment Retry - A payment request fails with
503 Service Unavailable, but an idempotency key guarantees that retrying cannot create a duplicate payment. -
Unsafe Payment Retry - A payment request returns
503 Service UnavailablewithRetry-After, but there is no idempotency key, so retrying could create a duplicate charge. -
Idempotent PUT - A
PUTrequest is sent successfully, but the connection is lost before the response arrives. -
Rate-Limited Order - An order request is rate-limited before processing and returns
429 Too Many RequestswithRetry-After. -
Payment Details Conflict - A multi-step payment flow reaches
/payments/details, which returns409 Conflictwithtransient-error: false. -
Non-Idempotent PATCH - A
PATCHrequest increases inventory, but the connection is lost before the client knows whether the change was already applied. - Idempotent DELETE - A session deletion request is sent, but the connection is lost before the response arrives.
-
Service Unavailable - A
GETrequest receives503 Service Unavailabletogether with aRetry-Afterheader. -
Stale Resource Version - A resource is updated by another client, so a
PATCHusing an old ETag fails with412 Precondition Failed. -
Rate Limit Without Retry-After - A safe
GETrequest repeatedly receives429 Too Many Requests, but the server does not provide aRetry-Afterheader. -
I'm a Teapot - A coffee request receives the legendary
418 I'm a Teapotresponse. -
Eventual Consistency - A newly created resource immediately returns
404 Not Foundwhen another service tries to use it. -
Expired Filter Workflow - A temporary filter times out several times and eventually returns
404 Not Found, so the original filter can no longer be used. -
Cached Cart Options - A request for cart options fails with
504 Gateway Timeout, but a recently expired cached response is still available throughstale-if-error.
Models Tested
Picking the models was pretty straightforward. First, I picked models I use a lot while working: GPT-5.6 Luna and GPT-5.6 Sol. From my experience, Luna is very efficient. It makes more mistakes than Sol, but it also uses significantly fewer tokens.
Then I added Gemini 3.7 Flash and Gemini 3.1 Pro Preview. I’ve used both quite a lot for text generation and CSS styling.
I also included Claude Sonnet 5 and Claude Haiku 4.5. I used Claude a lot for coding before, but these days I’ve mostly switched to GPT-5.6 because it feels more efficient for my workflow.
Lastly, I added GLM-5 and DeepSeek-R1 mostly out of curiosity to see how well they would perform.
I also tried to include Grok, but I immediately hit 404 Not Found errors with both Grok 4.5 and Grok 4.6, so they are not included in the results.
The goal wasn’t to test every available model, but to compare a mix of models I already use with a few I was curious about.
Findings
Okay, first let's have a look at the leaderboard and the results from the first run.
We can see that the best model, with a score of 1.00, was GPT-5.6 Sol, followed by Gemini 3.7 Flash with one mistake. Then came four models with two mistakes each, DeepSeek-R1 with three, and Claude Haiku 4.5 finished last with five.
When we look closer, we can see that most of the models made a mistake in the Unsafe Payment Retry task. The task itself is tricky because of the Retry-After header, but repeating that request could result in a duplicate charge. Only GPT-5.6 Sol and Gemini 3.1 Pro Preview answered correctly. The rest of the models returned something similar to:
Decision: YES_AFTER_DELAY
Reason: The server explicitly requests waiting 10 seconds before attempting the request again.
So most of the models fell into the trap of following the obvious HTTP signal, even though it conflicted with the wider context. Retry-After says “retry,” but a payment without idempotency says “maybe don't.” And to be honest, I am glad that most of the models failed. A small win for the author, who can still trick AI sometimes. 😄
I was also surprised that some models failed on tasks with idempotent HTTP methods: Idempotent PUT and Idempotent DELETE. But they did not fail completely. Most of them answered YES_AFTER_DELAY because they assumed a temporary network issue that could be resolved after a delay.
These YES versus YES_AFTER_DELAY cases show that a model can understand that retrying is safe but disagree about when to retry. That is different from misunderstanding the scenario. For version 2, I could therefore allow multiple valid Decision values for some scenarios. So in these cases, AI actually trained me a little.
Three Runs, Not Just One
That was only the first run, but I didn't want to build the whole experiment around a single run, so I ran the same benchmark two more times to check model stability and collect more data.
| Model | Run 1 | Run 2 | Run 3 | Avg. |
|---|---|---|---|---|
| GPT-5.6 Sol | 14/14 | 14/14 | 13/14 | 13.67 |
| Gemini 3.7 Flash | 13/14 | 13/14 | 13/14 | 13.00 |
| Gemini 3.1 Pro Preview | 12/14 | 12/14 | 13/14 | 12.33 |
| GLM-5 | 12/14 | 13/14 | 12/14 | 12.33 |
| Claude Sonnet 5 | 12/14 | 12/14 | 12/14 | 12.00 |
| GPT-5.6 Luna | 12/14 | 12/14 | 12/14 | 12.00 |
| DeepSeek-R1 | 11/14 | 10/14 | 10/14 | 10.33 |
| Claude Haiku 4.5 | 9/14 | 9/14 | 9/14 | 9.00 |
Overall, the results were actually quite stable. Out of the 112 model-scenario combinations (8 models × 14 scenarios), 102 had the same correct/incorrect outcome in all three runs, which is about 91%. But the interesting part is that the same score did not always mean the same behavior. Some models were stable in the number of correct decisions while changing which scenarios they got wrong. A score of 12/14 in all three runs does not necessarily mean that the model made the same decisions every time.
This was especially visible in Expired Filter Workflow, where four models changed their result between runs: GPT-5.6 Sol, Gemini 3.1 Pro Preview, GLM-5, and DeepSeek-R1. Even GPT-5.6 Sol, which scored 14/14 in the first two runs, changed its answer in the third run from YES to YES_AFTER_DELAY. It still decided that the request should be retried, but disagreed about when. This is exactly the type of scenario that made me question whether exact YES versus YES_AFTER_DELAY scoring is always the right approach.
The additional runs also made the Unsafe Payment Retry finding much stronger. It was answered correctly only 6 times out of 24 attempts, just 25%. Even more interestingly, the result was completely consistent across runs. GPT-5.6 Sol and Gemini 3.1 Pro Preview got it right all three times, while the other six models missed it in every run. So this looks less like random model variation and more like a systematic trap in how the models interpreted the scenario.
On the other hand, seven of the 14 scenarios were answered correctly in all 24 attempts. The straightforward retry cases were therefore not really what separated the models. The differences started to appear when retry timing, wider context, or multi-step state became important.
Model Efficiency
Now let's have a look at efficiency. I will use the graph from the third run because the models ended up in roughly the same areas of the chart across all three runs.
GPT-5.6 Luna seems to be the most efficient model, and I am not surprised. I mostly use it at work because it is cheap and usually gets me close to the final solution.
Claude Haiku 4.5 is also in the efficient corner, but it had the worst results of all the tested models and was still more expensive than GPT-5.6 Luna.
The rest of the models are in the upper-right corner, and I was really surprised that DeepSeek-R1 ended up as the second most expensive model. So I looked at its output and found that DeepSeek-R1 answered with anywhere from 4,000 to 9,000 characters, even though it was supposed to answer only with Decision and Reason. All of the other models followed that instruction, so in this case the cost difference wasn't only about token pricing, but also about how well the model followed the requested output format.
What I Learned
When I designed the experiment, I was actually worried that the scenarios would be too easy for today's AI models. The results showed that this is not always the case. Yes, I could definitely improve the benchmark, especially the ambiguity between YES and YES_AFTER_DELAY. But even the strongest models still made mistakes once the decision depended on more than just the obvious HTTP signal. Retry decisions in real systems are not always simple. AI can definitely help us reason through the difficult cases, but the wider context still matters, and blindly following the model's answer can be risky.
My Benchmark
To Retry or Not to Retry? Benchmark
Kaggle was completely new to me and, to be honest, I am glad that I discovered it through the DEV.to challenge. I expected the setup to be much more complicated, but creating this small experimental benchmark directly in the UI was surprisingly straightforward. I will definitely explore Kaggle more, especially after seeing how much there is to explore beyond traditional ML competitions.


Top comments (8)
Wow, what a great article!!! Not only is it incredibly practical for engineers and architects, but it also highlights something bigger: nope, we still can't blindly trust coding agents, no matter what some people try to convince us of. 😅
Now I'm seriously tempted to send this to a few self-proclaimed "experts". 🤣
Oh, thank you! Glad you liked it, Sylwia.
Yep, totally agree. Kaggle doesn't allow all the latest models, but even the best ones can make mistakes, so a good code review is still necessary.
I think the benchmark has some practical value, but this first version still has its limitations. The YES vs YES_AFTER_DELAY decisions are debatable, and more runs would help too. So I'm sure those self-proclaimed experts would find something to argue about. 🤣
Btw, how was Frontkon? Did you become even more famous? 😁
Hahahaha, I have no idea if I'm any more famous now! 🤣 But apparently, the head of the conference has appointed me Frontkon's ambassador to Poland, and my mission for next year is to bring 60 Polish girls with me! 🤣🤣🤣
The Unsafe Payment Retry row is the most informative one, but check how its label is defined before reading it as a model error. A 503 usually comes from a load balancer or an overloaded tier that never ran the handler, so YES_AFTER_DELAY with a Retry-After header is a defensible answer; NO is the conservative policy answer because a 503 doesn't prove the handler didn't run. The Reason lines are where you can see which models knew that and which just followed the header.
I'd split the scoring in two. One column for "would this have caused a duplicate effect" (the NO class, where a miss is costly), and one for YES vs YES_AFTER_DELAY, which is only about timing. Right now a model that says YES on the idempotent DELETE and one that says YES_AFTER_DELAY on an unsafe POST lose the same single point, though only one of them could charge a customer twice.
On the leaderboard: with 14 scenarios, one item is 7 points. The 95% Wilson interval for 14/14 is roughly 78-100%, for 12/14 roughly 60-96%, and even Haiku's 9/14 is about 39-84%. So the ordering between models is not separated by this run; the per-scenario pattern (which items the models miss in common) is the real result. Running each scenario 5 times per model would also show whether the Retry-After miss is stable or a coin flip for the models that got it right.
Thanks for the detailed feedback. I agree that separating retry safety from retry timing would make the benchmark much more meaningful. It's something I started thinking about after seeing the YES vs YES_AFTER_DELAY results.
For the payment scenario, I still think NO is the right answer when there's no idempotency guarantee. A 503 doesn't prove the payment was processed, but it doesn't prove the opposite either, and that's exactly the risk I wanted to test.
I actually ran the benchmark three times and interestingly, the Unsafe Payment Retry results were completely consistent across all three runs. More runs would definitely be better, but this is still a small experiment anyway.
Definitely some good ideas here for a potential v2!
Agreed on NO for the payment case: with no idempotency key, a 503 is ambiguous about whether the charge went through, so a blind retry can double-charge.
One caution on the three identical runs: if the settings were the same, matching results show the model is stable on that prompt, not that the answer is right. That is still useful for the payment item, since a stable NO is what you want there. For v2, a variant of the same scenario with an
Idempotency-Keyheader present would show whether the model flips to YES only when it should.daniel, this is a brilliant angle for an ai benchmark. testing models on contextual reasoning (idempotency, http methods, state) rather than just raw code generation is exactly what the industry needs right now.
your "unsafe payment retry" trap is the perfect example of why blind ai automation is dangerous. as someone building an ai coding ecosystem (koda), this is my biggest fear: an agent confidently executing a non-idempotent mutation just because it saw a
retry-afterheader, completely ignoring the lack of an idempotency key.i also completely agree with @arhancanli's point about splitting the scoring. a model guessing the wrong timing (
yesvsyes_after_delay) is a minor flaw, but a model causing a duplicate charge is a catastrophic failure. weighting safety heavily over timing would make this benchmark even more powerful for v2.thanks for proving that good old-fashioned architectural principles (like idempotency and deterministic fallbacks) still trump raw model size! 🐯
Hmm, why are they still using DeepSeek-R1? It's a bit outdated. They should try the latest DeepSeek-V4.1 Flash.
and Good Luck!