This is a submission for the Kaggle Benchmarking Challenge
Those who read my articles know that a lot of them are actually benchmarks of something. Mostly .NET related, but still, someone could say that this challenge should be pretty close to what I usually do.
The opposite was true.
Building a code benchmark and benchmarking AI models are two different worlds, so when I first saw this challenge, I had no idea what exactly I should benchmark. Comparing models on coding tasks felt too generic, and I didn't want to create a benchmark just for the sake of having one.
Then I looked at the topics of my last couple of articles. A lot of them were about APIs, resilience, failures, and how systems behave when something goes wrong. And that gave me an experiment idea: What if I benchmark AI models on one very simple question: To Retry or Not to Retry?
A 503 Service Unavailable does not automatically mean that retrying is safe. A POST request may already have been processed. An idempotency key can completely change the answer. A timeout may happen before the server receives anything or after it has already changed some state.
So the HTTP status code alone is often not enough. The model has to understand the whole situation.
Top comments (0)