<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mzeeshan127</title>
    <description>The latest articles on DEV Community by Mzeeshan127 (@mzeeshan127).</description>
    <link>https://dev.to/mzeeshan127</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4157237%2F32fdac35-e3f7-4051-b4fa-02fcc1d1a803.png</url>
      <title>DEV Community: Mzeeshan127</title>
      <link>https://dev.to/mzeeshan127</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mzeeshan127"/>
    <language>en</language>
    <item>
      <title>Can GPT-5.4 mini handle least-privilege cloud incidents? A 16-case benchmark</title>
      <dc:creator>Mzeeshan127</dc:creator>
      <pubDate>Fri, 02 Oct 2026 12:20:51 +0000</pubDate>
      <link>https://dev.to/mzeeshan127/can-gpt-54-mini-handle-least-privilege-cloud-incidents-a-16-case-benchmark-428i</link>
      <guid>https://dev.to/mzeeshan127/can-gpt-54-mini-handle-least-privilege-cloud-incidents-a-16-case-benchmark-428i</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;I built a small decision benchmark for cloud operations and incident response. It presents 16 fully synthetic situations involving exposed credentials, access scope, suspicious accounts, evidence preservation, risky commands, storage exposure, firewall changes, and approval boundaries. Each case expects one documented action identifier; scoring is exact match. I was interested in whether models could choose a safe, proportionate action while respecting least privilege and approval boundaries.&lt;/p&gt;

&lt;p&gt;No cloud APIs, production infrastructure, real credentials, or customer data are used.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;I ran OpenAI GPT-5.4 mini (Kaggle model listing: GPT-5.4 mini) against the task. I started with a single model to establish a reproducible baseline; additional candidates were not successfully evaluated in this run, so this is not a model-to-model comparison.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;p&gt;GPT-5.4 mini scored &lt;strong&gt;75.0% (12/16)&lt;/strong&gt; exact-match accuracy on the fixed set, evaluated October 2, 2026. That means four choices did not match the reference action identifiers. The aggregate score does not show which case types caused those mismatches, so I cannot responsibly infer a specific weakness from this run. The next step is to inspect per-case outcomes and run a broader model lineup under the same task conditions.&lt;/p&gt;

&lt;p&gt;This is a small, fixed synthetic set, not evidence of real-world security competence, reasoning quality, calibration, or performance on live incidents.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.kaggle.com/benchmarks/zeeshan1271/least-privilege-cloud-operations" rel="noopener noreferrer"&gt;Least-Privilege Cloud Operations on Kaggle&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The public benchmark includes the task, scoring description, and recorded result.&lt;/p&gt;




&lt;p&gt;AI tools assisted with parts of preparing this benchmark and submission.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
