This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
I built a small decision benchmark for cloud operations and incident response. It presents 16 fully synthetic situations involving exposed credentials, access scope, suspicious accounts, evidence preservation, risky commands, storage exposure, firewall changes, and approval boundaries. Each case expects one documented action identifier; scoring is exact match. I was interested in whether models could choose a safe, proportionate action while respecting least privilege and approval boundaries.
No cloud APIs, production infrastructure, real credentials, or customer data are used.
Models Tested
I ran OpenAI GPT-5.4 mini (Kaggle model listing: GPT-5.4 mini) against the task. I started with a single model to establish a reproducible baseline; additional candidates were not successfully evaluated in this run, so this is not a model-to-model comparison.
Findings
GPT-5.4 mini scored 75.0% (12/16) exact-match accuracy on the fixed set, evaluated October 2, 2026. That means four choices did not match the reference action identifiers. The aggregate score does not show which case types caused those mismatches, so I cannot responsibly infer a specific weakness from this run. The next step is to inspect per-case outcomes and run a broader model lineup under the same task conditions.
This is a small, fixed synthetic set, not evidence of real-world security competence, reasoning quality, calibration, or performance on live incidents.
My Benchmark
Least-Privilege Cloud Operations on Kaggle
The public benchmark includes the task, scoring description, and recorded result.
AI tools assisted with parts of preparing this benchmark and submission.
Top comments (1)
Good call to say plainly that one model and 16 cases isn't a comparison yet. Two things would make the next run say much more.
On 16 items, 12/16 has a 95% interval of roughly 50% to 90%, so two models' overall scores will overlap unless the gap is huge. When you add models, compare them case by case on the same 16 (count the cases where one is right and the other wrong) rather than by their totals; that paired view is far more sensitive. Running every case several times at the model's normal temperature also shows which of the four misses are stable and which are coin flips, and that per-case table is what lets you name a weakness.
The other check I'd add before more models: a trivial policy baseline. Score "always pick the most cautious action" (escalate / request approval / preserve evidence, whichever your action IDs include) on the same 16 cases. In incident-response sets the reference answer is often the conservative one, so a never-take-risk rule can score surprisingly well. If it gets 9 or 10 of 16, the model's 12 means much less than if it gets 4.