This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
Reliability under constraints: when a model must satisfy several rules at once, and the prompt is padded with distractors or conflicting notes, which rules slip?
ModelBench has 16 tasks in 8 control/stress pairs across four families (constraints, structured output, distractors, compound tasks). A control is a plain task; its stress version adds extra constraints, conflicting instructions or irrelevant facts. Every answer is scored by deterministic code (JSON parsing, word limits, exact values, forbidden words), never by an AI judge. A task's score is the share of its checks passed. Infrastructure errors are never counted against a model.
Models Tested
Five models on Kaggle, one run each: gpt-oss-20b, Gemini 3.7 Flash, Claude Sonnet 5.5, Gemma 4 31B and DeepSeek-R1.
Findings
1. A huge spread on identical prompts. gpt-oss-20b scored 96.7%, Gemini 3.7 Flash 94%, Claude Sonnet 5.5 84%, Gemma 4 31B 68% and DeepSeek-R1 16%.
2. Score did not follow cost. On Kaggle's score-vs-cost chart, gpt-oss-20b sits in the "efficient" corner, while DeepSeek-R1 is among the most expensive points with the lowest score.
3. Strong models still slip on format discipline. In an earlier manual pilot, Claude Sonnet 5.5 passed all 8 original tasks, but on a harder set it gave a correct multi-step answer (24.19) while showing its working, even though the prompt said "reply with only the number". Gemini 3.6 Flash passed 7 of 8 harder tasks and dropped a task from a list after a distractor said sprint planning had moved (one run, so only a hypothesis).
4. What surprised me. Claude scored 100% on the first 8 tasks in my manual pilot but 84% on all 16 on Kaggle. I haven't diagnosed why. The 8 added tasks are harder, but one run per model is a weak basis for a conclusion.
My Benchmark
https://www.kaggle.com/benchmarks/ezeprincejr/modelbench-reliability-under-constraints
Limitations
Only 16 tasks and one run per model, so small differences are not reliable. Scoring is equal-weight and provisional. Eight of the tasks were written after a two-model pilot, so they are not blind. A JSON reply wrapped in a code fence counts as a format failure. I did not diagnose why DeepSeek-R1 scored so low (it could be a real format problem or my scoring being too strict). I have not split scores into control vs stress per model. Next I'd add repeated runs, more models and a per-task failure analysis.
Prototype dashboard (optional)
I also built a companion dashboard as an installable web app: https://modelbench-ai-c14be.web.app
It is only a prototype. The official, authoritative results are the ones on the Kaggle leaderboard above. The dashboard shows an "awaiting results" screen until results are loaded into it, and a labelled demo-data mode for previewing the layout.
Top comments (2)
The eight control/stress pairs are the most useful structure here, and I would report their paired changes before the overall ranking. A model can earn a high average while losing exactly the requirement introduced in a stress prompt.
Because tasks have different numbers of checks, a share-of-checks score can also dilute one consequential failure among several easy formatting passes. Keeping an all-required-checks pass rate beside the fractional score would expose that. The 24.19 answer with extra working is a good concrete example: numerical correctness and exact output-contract compliance should remain separately visible.
Official Platform Update
Security protocols have been updated for all developer accounts.