This is a submission for the Kaggle Benchmarking Challenge.
What I Benchmarked
Hindsight asks a simple question: if you had shown today's models the code behind history's most expensive software failures, the day before it shipped, would they have caught the bug?
I spent the last few weeks reading accident reports for a series of short animated explainers about famous failures: the Ariane 5 inertial reference software, the Patriot battery at Dhahran, Knight Capital's 45 minutes, the CrowdStrike channel file, the spreadsheet behind Reinhart and Rogoff. Every report ends the same way: the defect was small, local and visible in a few lines, and nobody saw it in time. That made me curious whether a code review by a model would have changed anything.
The benchmark has 26 cases. Each is a documented failure with a published root cause (an inquiry board, a GAO or SEC report, a vendor post-mortem or a peer-reviewed paper), rebuilt as a minimal artifact: a function, a shell script, an Excel formula, an RTOS configuration. No names, no dates. The model gets the artifact and one line about the system, and answers in a fixed format: VERDICT: DEFECT / NO_DEFECT, what fails, the trigger, the consequence. A judge model compares the answer with the published root cause.
Every case comes in four forms, which are the four tasks of the benchmark:
| Task | What the model sees | What it measures |
|---|---|---|
| Original | the reconstruction in its original setting (Ada on a launcher, Pascal on a radiotherapy machine) | can it catch the failure? |
| Disguised | the same mechanism moved to another domain, language and naming (the Ariane conversion becomes a Python drone telemetry packer) | is it reasoning, or recognising a famous story? |
| Disguised + context | the disguised artifact plus the one operating fact that triggered the real failure ("the new drone flies five times faster") | how much does knowing the environment help? |
| Fixed | the original artifact after the fix that was applied | does it raise false alarms on code that is now correct? |
The gap between Original and Disguised is a memorisation signal. The gap between Disguised and Disguised + context measures something the accident reports keep returning to: much of this code was correct for the system it was written for, and failed in a new environment.
One case, four forms
Here is the Ariane 5 case. The original form is close to what flew in 1996: a 64-bit float converted to a 16-bit integer, protected on one variable and not on the other, because analysis of the old rocket's trajectory showed it stayed small.
procedure Update_Alignment is
HB_Bus : Integer_16; -- horizontal bias, 16-bit field on the data bus
VB_Bus : Integer_16; -- vertical bias
begin
Horizontal_Bias := Compute_Horizontal_Bias (Horizontal_Velocity); -- Long_Float
Vertical_Bias := Compute_Vertical_Bias (Vertical_Velocity); -- Long_Float
VB_Bus := Integer_16 (Saturate (Vertical_Bias, -32_768.0, 32_767.0));
HB_Bus := Integer_16 (Horizontal_Bias);
-- no saturation on HB: analysis of the trajectory showed it stays small
Bus.Put (Field => VB, Value => VB_Bus);
Bus.Put (Field => HB, Value => HB_Bus);
end Update_Alignment; -- no exception handler in this procedure
The disguised form keeps the mechanism and changes everything a model could pattern-match on: a cargo drone, Python, a radio packet, millimetres instead of velocities.
def pack_drift(nav):
# radio link fields are signed 16-bit integers (struct code 'h')
alt_drift = clamp(int(nav.vertical_drift_mm), -32768, 32767)
lat_drift = int(nav.lateral_drift_mm) # analysis showed lateral drift stays small
return struct.pack('<hh', alt_drift, lat_drift)
def flight_loop(nav, radio):
while True:
nav.update()
radio.send(pack_drift(nav))
The context form adds the one sentence the Ariane 5 team did not connect to this code: "the new drone cruises about five times faster than the quadcopter, so its lateral drift estimate during the climb reaches 40,000 to 60,000 mm." And the fixed form is the Ada original with Horizontal_Bias saturated as well, so a model that still says "this will overflow" is now wrong.
How an answer is scored
A case counts as caught only if the model says DEFECT and the judge confirms that the answer names the same mechanism and trigger as the published root cause. "This conversion could overflow" is not enough for Ariane 5; the answer has to say that the horizontal bias goes past 32,767 and that the unhandled exception stops the unit. A correct verdict with the wrong reason is a miss.
On the fixed versions the rule flips: an answer passes unless it claims that the historical failure will still happen. Raising some other concern is allowed, and I report separately how often models flag something on correct code.
The judge is Kaggle's default evaluation model. To check it, I re-read 64 answers by hand: the 44 disguised answers from the first nine models that said DEFECT but were marked as misses, plus 20 random passes. I agreed with the judge on 61. All three disagreements were the judge being too strict, so if anything the scores below are slightly low.
Models Tested
Thirteen models from six vendors, large and small. Ranking them was half the point; the other half was seeing whether size and vendor change what kind of bug gets missed:
- Anthropic: Claude Opus 5.5, Claude Sonnet 5.5, Claude Haiku 4.5
- OpenAI: GPT-6.1 Sol, GPT-5.4 mini, gpt-oss-120b (open weights)
- Google: Gemini 3.1 Pro (preview), Gemini 3.8 Flash, Gemini 3.7 Flash, Gemma 4 31B (open weights)
- xAI: Grok 4.20 (reasoning)
- DeepSeek-R1 and Qwen3 Coder 480B (open weights)
Every model gets the same prompt with the platform's default sampling settings and a 4,000-token output cap, one case per fresh chat.
Findings
The famous bugs are solved, even in disguise
Eight cases were caught by all thirteen models in their disguised form: Ariane 5, the Patriot clock, Mars Climate Orbiter, the Boeing 787 counter, the Azure leap-day certificate, the AT&T switch crash, CrowdStrike and the Reinhart-Rogoff spreadsheet. Ten more were caught by at least ten of the thirteen. A Python drone packer with an unclamped 16-bit field does not fool anyone any more. These are bugs you can see by reading the code once: a narrowing conversion, a counter that wraps, a quantity in the wrong unit, a range that stops a few rows short.
What they miss only shows up when the code runs
Seven cases were caught by eight of the thirteen models or fewer:
| Case | Caught | What you have to imagine |
|---|---|---|
| Mars Polar Lander, 1999 | 46% | a vibration spike during leg deployment, latched as "touchdown" long before the ground |
| Vancouver Stock Exchange index, 1983 | 46% | a truncation to three decimals, repeated on every trade for 22 months |
| Cloudflare WAF, 2019 | 46% | one regex meeting a long input and backtracking until the CPU is gone |
| England's COVID test results, 2020 | 46% | an export format that stops at 65,536 rows |
| Amazon S3, 2017 | 54% | a maintenance command removing far more capacity than the operator meant |
| Schiaparelli, 2016 | 54% | an inertial unit that saturates for one second while the parachute opens |
| London Whale, 2012 | 62% | a volatility that is divided by a sum where it should be an average |
None of these is hard to read. Each one needs you to run the code in your head under a condition that is not on the page: a long input, a vibration, the 22nd month, the 65,537th row. That is exactly what the accident reports say the engineers failed to do, and it is the part the models still struggle with.
Right verdict, wrong reason
Across the 338 disguised answers the models said DEFECT 97% of the time. Of the 63 misses, 54 still said DEFECT: they found a problem, just not the one that brought the system down. Some of the wrong reasons are reasonable review comments. Some are invented.
- On the Cloudflare regex, Claude Haiku 4.5 warned about SQL injection ("
DROP TABLE users;--"). The regex never touches a database; the problem is that it can take exponential time. - On the Vancouver index, Gemini 3.1 Pro blamed floating point: "
0.58 * 100evaluates to57.99999999999999". True, and a one-cent error, while truncating on every update took half the value of the index. - On England's spreadsheet export, Gemini 3.8 Flash said the export call "accepts only a single argument" and will crash. It takes a file path just fine, and the answer never mentions the row limit of the
.xlsformat. - On Mars Polar Lander, GPT-6.1 Sol found the latch that is never cleared, but blamed a previous trip instead of the vibration on the way down: the right line of code with the wrong story.
This is why a DEFECT verdict on its own means very little here, and why the benchmark scores the mechanism and the trigger, not the verdict.
My reconstructions had bugs of their own
Reading the misses by hand, I found that three of my artifacts contained a second, unintended defect, and some models reported that one instead of the historical one: the Cloudflare regex also rejects a valid single comparison, the London Whale sheet divides by zero when both weeks are empty, and the S3 drain script deletes nodes without waiting for the evictions to finish. Those answers are correct review comments and still count as misses, because they are not what happened. When you rebuild a failure, you rebuild more than one failure. I have left the cases as they were run and listed the extra defects in the benchmark notes.
Memory, context and false alarms
Averaged over all thirteen models, the four tasks score 89% (original), 81% (disguised), 89% (disguised with context) and 96% (fixed version, no false alarm).
Memory helps less than I expected, and not where I expected. The memorisation gap, original minus disguised, is eight points on average. It is largest for DeepSeek-R1 (19 points), gpt-oss-120b and Gemma 4 31B (15 each), and zero or below for Claude Sonnet 5.5, Claude Opus 5.5 and Qwen3 Coder. I also counted how often an answer to the original form names the real incident ("this is the Ariane 5 failure"). Opus did that on 81% of cases, Gemini 3.1 Pro on 58%, Sonnet on 46%, and nearly every other model on 8% or less. The models that recognise the stories most often are the ones that lose nothing when the story is taken away, so recognising a case and depending on that recognition are two different things. The gap belongs to the mid-sized and open models.
One sentence of context is worth as much as the original setting. Adding the operating fact to the disguised artifact brings the average from 81% back to 89%, level with the original. The cases it rescues are exactly the hard ones from the table above: the spreadsheet export (7 models went from miss to catch), the regex, the lander and the parachute IMU (6 each). DeepSeek-R1 gains 23 points, GPT-5.4 mini and Gemma 15. Grok 4.20 is the exception: it lost three cases when given the context (Zune, S3 and the London Whale sheet), and the Zune case lost five models in total, because the hint about the last day of a leap year pulled answers toward the date arithmetic and away from the loop that never ends.
Fixed code is rarely condemned, but often criticised. Only 15 of 338 answers claimed the historical failure would still happen on the fixed version, and six of those were the same case: Schiaparelli, where the fix drops the one-second hold on the saturated rate and adds a plausibility check, and six models still argued that the altitude could go negative. But models flagged some defect on 44% of the fixed artifacts. Claude Haiku 4.5 did so on 96% of them, GPT-5.4 mini on 88%, gpt-oss-120b on 81%; Sonnet 5.5 never did and Opus 5.5 once. For a code reviewer that difference matters as much as the catch rate: a tool that objects to everything teaches people to stop reading its comments.
My Benchmark
The benchmark on Kaggle: Hindsight: Catching History's Costliest Bugs, with the leaderboard for all thirteen models.
The four tasks, each with the reconstructions, the sources for every case and the per-model runs:


Top comments (0)