<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Eze Prince </title>
    <description>The latest articles on DEV Community by Eze Prince  (@eze_princejr_1dcc53250ce).</description>
    <link>https://dev.to/eze_princejr_1dcc53250ce</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4170080%2F370c022f-daa5-4189-a212-29b6769e65d1.webp</url>
      <title>DEV Community: Eze Prince </title>
      <link>https://dev.to/eze_princejr_1dcc53250ce</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/eze_princejr_1dcc53250ce"/>
    <language>en</language>
    <item>
      <title>Do AI models still follow instructions when the rules stack up? A control-vs-stress benchmark</title>
      <dc:creator>Eze Prince </dc:creator>
      <pubDate>Sun, 11 Oct 2026 07:52:22 +0000</pubDate>
      <link>https://dev.to/eze_princejr_1dcc53250ce/do-ai-models-still-follow-instructions-when-the-rules-stack-up-a-control-vs-stress-benchmark-1bkf</link>
      <guid>https://dev.to/eze_princejr_1dcc53250ce/do-ai-models-still-follow-instructions-when-the-rules-stack-up-a-control-vs-stress-benchmark-1bkf</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;Reliability under constraints: when a model must satisfy several rules at once, and the prompt is padded with distractors or conflicting notes, which rules slip?&lt;/p&gt;

&lt;p&gt;ModelBench has 16 tasks in 8 control/stress pairs across four families (constraints, structured output, distractors, compound tasks). A control is a plain task; its stress version adds extra constraints, conflicting instructions or irrelevant facts. Every answer is scored by deterministic code (JSON parsing, word limits, exact values, forbidden words), never by an AI judge. A task's score is the share of its checks passed. Infrastructure errors are never counted against a model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;Five models on Kaggle, one run each: gpt-oss-20b, Gemini 3.7 Flash, Claude Sonnet 5.5, Gemma 4 31B and DeepSeek-R1.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. A huge spread on identical prompts.&lt;/strong&gt; gpt-oss-20b scored 96.7%, Gemini 3.7 Flash 94%, Claude Sonnet 5.5 84%, Gemma 4 31B 68% and DeepSeek-R1 16%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Score did not follow cost.&lt;/strong&gt; On Kaggle's score-vs-cost chart, gpt-oss-20b sits in the "efficient" corner, while DeepSeek-R1 is among the most expensive points with the lowest score.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Strong models still slip on format discipline.&lt;/strong&gt; In an earlier manual pilot, Claude Sonnet 5.5 passed all 8 original tasks, but on a harder set it gave a correct multi-step answer (24.19) while showing its working, even though the prompt said "reply with only the number". Gemini 3.6 Flash passed 7 of 8 harder tasks and dropped a task from a list after a distractor said sprint planning had moved (one run, so only a hypothesis).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. What surprised me.&lt;/strong&gt; Claude scored 100% on the first 8 tasks in my manual pilot but 84% on all 16 on Kaggle. I haven't diagnosed why. The 8 added tasks are harder, but one run per model is a weak basis for a conclusion.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.kaggle.com/benchmarks/ezeprincejr/modelbench-reliability-under-constraints" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/ezeprincejr/modelbench-reliability-under-constraints&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Limitations
&lt;/h2&gt;

&lt;p&gt;Only 16 tasks and one run per model, so small differences are not reliable. Scoring is equal-weight and provisional. Eight of the tasks were written after a two-model pilot, so they are not blind. A JSON reply wrapped in a code fence counts as a format failure. I did not diagnose why DeepSeek-R1 scored so low (it could be a real format problem or my scoring being too strict). I have not split scores into control vs stress per model. Next I'd add repeated runs, more models and a per-task failure analysis.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prototype dashboard (optional)
&lt;/h2&gt;

&lt;p&gt;I also built a companion dashboard as an installable web app: &lt;a href="https://modelbench-ai-c14be.web.app" rel="noopener noreferrer"&gt;https://modelbench-ai-c14be.web.app&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It is only a prototype. The official, authoritative results are the ones on the Kaggle leaderboard above. The dashboard shows an "awaiting results" screen until results are loaded into it, and a labelled demo-data mode for previewing the layout.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>INVORA PRO business management and POS system... now with APIs...</title>
      <dc:creator>Eze Prince </dc:creator>
      <pubDate>Thu, 08 Oct 2026 10:16:54 +0000</pubDate>
      <link>https://dev.to/eze_princejr_1dcc53250ce/invora-pro-business-management-and-pos-system-now-with-apis-2h9h</link>
      <guid>https://dev.to/eze_princejr_1dcc53250ce/invora-pro-business-management-and-pos-system-now-with-apis-2h9h</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjieoy7ybkj7uqrq0mlma.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjieoy7ybkj7uqrq0mlma.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;br&gt;
&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmxbukcj4y8v55gf5c40j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmxbukcj4y8v55gf5c40j.png" alt="INVORA PRO v3.3.2" width="800" height="1200"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
