DEV Community

Cover image for Do AI models still follow instructions when the rules stack up? A control-vs-stress benchmark
Eze Prince
Eze Prince

Posted on

Do AI models still follow instructions when the rules stack up? A control-vs-stress benchmark

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

Reliability under constraints: when a model must satisfy several rules at once, and the prompt is padded with distractors or conflicting notes, which rules slip?

ModelBench has 16 tasks in 8 control/stress pairs across four families (constraints, structured output, distractors, compound tasks). A control is a plain task; its stress version adds extra constraints, conflicting instructions or irrelevant facts. Every answer is scored by deterministic code (JSON parsing, word limits, exact values, forbidden words), never by an AI judge. A task's score is the share of its checks passed. Infrastructure errors are never counted against a model.

Models Tested

Five models on Kaggle, one run each: gpt-oss-20b, Gemini 3.7 Flash, Claude Sonnet 5.5, Gemma 4 31B and DeepSeek-R1.

Findings

1. A huge spread on identical prompts. gpt-oss-20b scored 96.7%, Gemini 3.7 Flash 94%, Claude Sonnet 5.5 84%, Gemma 4 31B 68% and DeepSeek-R1 16%.

2. Score did not follow cost. On Kaggle's score-vs-cost chart, gpt-oss-20b sits in the "efficient" corner, while DeepSeek-R1 is among the most expensive points with the lowest score.

3. Strong models still slip on format discipline. In an earlier manual pilot, Claude Sonnet 5.5 passed all 8 original tasks, but on a harder set it gave a correct multi-step answer (24.19) while showing its working, even though the prompt said "reply with only the number". Gemini 3.6 Flash passed 7 of 8 harder tasks and dropped a task from a list after a distractor said sprint planning had moved (one run, so only a hypothesis).

4. What surprised me. Claude scored 100% on the first 8 tasks in my manual pilot but 84% on all 16 on Kaggle. I haven't diagnosed why. The 8 added tasks are harder, but one run per model is a weak basis for a conclusion.

My Benchmark

https://www.kaggle.com/benchmarks/ezeprincejr/modelbench-reliability-under-constraints

Limitations

Only 16 tasks and one run per model, so small differences are not reliable. Scoring is equal-weight and provisional. Eight of the tasks were written after a two-model pilot, so they are not blind. A JSON reply wrapped in a code fence counts as a format failure. I did not diagnose why DeepSeek-R1 scored so low (it could be a real format problem or my scoring being too strict). I have not split scores into control vs stress per model. Next I'd add repeated runs, more models and a per-task failure analysis.

Prototype dashboard (optional)

I also built a companion dashboard as an installable web app: https://modelbench-ai-c14be.web.app

It is only a prototype. The official, authoritative results are the ones on the Kaggle leaderboard above. The dashboard shows an "awaiting results" screen until results are loaded into it, and a labelled demo-data mode for previewing the layout.

Top comments (2)

Collapse
 
ahmetozel profile image
Ahmet Özel •

The eight control/stress pairs are the most useful structure here, and I would report their paired changes before the overall ranking. A model can earn a high average while losing exactly the requirement introduced in a stress prompt.

Because tasks have different numbers of checks, a share-of-checks score can also dilute one consequential failure among several easy formatting passes. Keeping an all-required-checks pass rate beside the fractional score would expose that. The 24.19 answer with extra working is a good concrete example: numerical correctness and exact output-contract compliance should remain separately visible.

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

Official Platform Update

Security protocols have been updated for all developer accounts.

  • tr.ee/dev-to