<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Hossam Elshahaby</title>
    <description>The latest articles on DEV Community by Hossam Elshahaby (@helshahaby).</description>
    <link>https://dev.to/helshahaby</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3332650%2Fec928ead-45e6-476e-b76b-e053098e12e6.png</url>
      <title>DEV Community: Hossam Elshahaby</title>
      <link>https://dev.to/helshahaby</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/helshahaby"/>
    <language>en</language>
    <item>
      <title>Kaggle Benchmarking Challenge</title>
      <dc:creator>Hossam Elshahaby</dc:creator>
      <pubDate>Tue, 06 Oct 2026 09:03:44 +0000</pubDate>
      <link>https://dev.to/helshahaby/kaggle-benchmarking-challenge-2973</link>
      <guid>https://dev.to/helshahaby/kaggle-benchmarking-challenge-2973</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h1&gt;
  
  
  ConstraintBench: What Happens When an AI Has Too Many Instructions?
&lt;/h1&gt;

&lt;p&gt;Most AI benchmarks ask a familiar question:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Did the model get the answer right?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I wanted to ask a slightly different one:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What happens when getting the answer right isn't enough?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Real-world prompts rarely contain one clean instruction. An AI agent may need to solve a problem, return an exact structure, include required information, exclude sensitive information, respect ordering and length requirements, perform calculations, and obey higher-priority rules—all in the same response.&lt;/p&gt;

&lt;p&gt;A model can therefore produce an answer that looks correct while still failing the application around it.&lt;/p&gt;

&lt;p&gt;That is what I built &lt;strong&gt;ConstraintBench&lt;/strong&gt; to measure.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;ConstraintBench measures &lt;strong&gt;multi-constraint instruction following under specification pressure&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Instead of testing only whether a model knows the correct answer, each case combines several requirements that must be satisfied simultaneously.&lt;/p&gt;

&lt;p&gt;The benchmark contains &lt;strong&gt;24 deterministic cases across four task families&lt;/strong&gt;, with six cases in each family.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Structured Extraction
&lt;/h3&gt;

&lt;p&gt;These tasks require models to extract information while simultaneously satisfying requirements such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;exact output structure;&lt;/li&gt;
&lt;li&gt;mandatory fields;&lt;/li&gt;
&lt;li&gt;prohibited content;&lt;/li&gt;
&lt;li&gt;ordering;&lt;/li&gt;
&lt;li&gt;formatting;&lt;/li&gt;
&lt;li&gt;numerical conditions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A response can contain all of the right information and still fail if it cannot be consumed reliably by another system.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Reasoning Under Constraints
&lt;/h3&gt;

&lt;p&gt;These cases combine reasoning or numerical problems with additional output requirements.&lt;/p&gt;

&lt;p&gt;The model has to solve the underlying problem &lt;strong&gt;and&lt;/strong&gt; preserve the surrounding specification.&lt;/p&gt;

&lt;p&gt;This lets the benchmark distinguish between:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The model knew the answer."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;and:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The model completed the task."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those are not always the same thing.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Transformation &amp;amp; Editing
&lt;/h3&gt;

&lt;p&gt;These tasks ask models to transform content while preserving specific information and satisfying multiple editing requirements.&lt;/p&gt;

&lt;p&gt;This resembles many real AI workflows: summarization, rewriting, document processing, data preparation, and agent-generated content.&lt;/p&gt;

&lt;p&gt;The challenge is not simply producing fluent text. The model must change exactly what it was asked to change without accidentally changing or losing something else.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Priority / Safety Preservation
&lt;/h3&gt;

&lt;p&gt;The final family tests whether important requirements survive when they compete with other instructions.&lt;/p&gt;

&lt;p&gt;These cases include constraints involving privacy, prohibited information, conditional behavior, and higher-priority requirements.&lt;/p&gt;

&lt;p&gt;This category interested me especially because not every constraint has the same consequence.&lt;/p&gt;

&lt;p&gt;Dropping a formatting requirement is inconvenient.&lt;/p&gt;

&lt;p&gt;Dropping a privacy or safety requirement can be much more serious.&lt;/p&gt;




&lt;h2&gt;
  
  
  How ConstraintBench Works
&lt;/h2&gt;

&lt;p&gt;Each task family contains six cases with increasing specification pressure.&lt;/p&gt;

&lt;p&gt;The cases combine constraints such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;produce valid structured output;&lt;/li&gt;
&lt;li&gt;include required information;&lt;/li&gt;
&lt;li&gt;exclude prohibited information;&lt;/li&gt;
&lt;li&gt;use an exact number of items;&lt;/li&gt;
&lt;li&gt;preserve a specified ordering;&lt;/li&gt;
&lt;li&gt;respect length requirements;&lt;/li&gt;
&lt;li&gt;perform numerical reasoning;&lt;/li&gt;
&lt;li&gt;follow conditional instructions;&lt;/li&gt;
&lt;li&gt;preserve higher-priority requirements.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;ConstraintBench does not use another language model as the judge.&lt;/p&gt;

&lt;p&gt;Instead, the cases use &lt;strong&gt;deterministic checks&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A required key exists or it does not.&lt;/p&gt;

&lt;p&gt;A prohibited value appears or it does not.&lt;/p&gt;

&lt;p&gt;A numerical condition is satisfied or it is not.&lt;/p&gt;

&lt;p&gt;A requested structure is valid or it is not.&lt;/p&gt;

&lt;p&gt;That makes failures easier to reproduce and interpret.&lt;/p&gt;

&lt;p&gt;The central question behind the benchmark is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What does a model forget first when it has too many things it cannot forget?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;I deliberately tested models from several different families rather than comparing only closely related frontier models.&lt;/p&gt;

&lt;p&gt;My Kaggle benchmark includes models from the Gemini, Gemma, Claude, GPT, Grok, GLM, DeepSeek, and Qwen families.&lt;/p&gt;

&lt;p&gt;The lineup includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Gemini 3.7 Flash&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gemma 4 26B A4B&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Grok 4.20 Reasoning&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Claude Haiku 4.5&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Claude Sonnet 4.6&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Claude Opus 4.6&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GPT-5.4&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;GLM-5&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;DeepSeek-R1&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Qwen 3 235B A22B&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I wanted this diversity because ConstraintBench is not intended to be another test of raw model intelligence.&lt;/p&gt;

&lt;p&gt;A smaller or faster model might be extremely dependable at structured compliance.&lt;/p&gt;

&lt;p&gt;A much larger reasoning model might solve a difficult problem correctly but miss an apparently minor output requirement.&lt;/p&gt;

&lt;p&gt;For agents and automated workflows, those differences can matter.&lt;/p&gt;




&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;p&gt;The most interesting result was not simply which model appeared at the top of the leaderboard.&lt;/p&gt;

&lt;p&gt;It was how differently models behaved across the four task families.&lt;/p&gt;

&lt;h3&gt;
  
  
  Correctness and compliance are different capabilities
&lt;/h3&gt;

&lt;p&gt;ConstraintBench repeatedly highlights a distinction that conventional accuracy metrics can hide:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A correct answer can still be a failed response.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Imagine an agent correctly calculates a transaction amount but produces invalid JSON.&lt;/p&gt;

&lt;p&gt;The reasoning was correct.&lt;/p&gt;

&lt;p&gt;The workflow still breaks.&lt;/p&gt;

&lt;p&gt;Or imagine a model produces an excellent summary but includes a piece of information it was explicitly instructed to omit.&lt;/p&gt;

&lt;p&gt;The text may look good.&lt;/p&gt;

&lt;p&gt;The specification was still violated.&lt;/p&gt;

&lt;p&gt;ConstraintBench treats those requirements as part of correctness rather than as optional presentation details.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reliability can be task-specific
&lt;/h3&gt;

&lt;p&gt;Another important observation was that performance could change sharply between categories.&lt;/p&gt;

&lt;p&gt;A model that handles priority preservation successfully does not necessarily show the same reliability on reasoning-under-constraints or transformation tasks.&lt;/p&gt;

&lt;p&gt;Likewise, strong reasoning ability does not automatically imply perfect structured-output compliance.&lt;/p&gt;

&lt;p&gt;This suggests that a single overall model score can hide something important.&lt;/p&gt;

&lt;p&gt;Models have something closer to a &lt;strong&gt;constraint-reliability profile&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For a production application, the shape of that profile may matter more than a small difference in an aggregate benchmark score.&lt;/p&gt;

&lt;h3&gt;
  
  
  Models do not necessarily fail gracefully
&lt;/h3&gt;

&lt;p&gt;I originally expected degradation to be relatively smooth.&lt;/p&gt;

&lt;p&gt;A model might satisfy almost everything under light specification pressure, then gradually miss more requirements as prompts become more complicated.&lt;/p&gt;

&lt;p&gt;The benchmark results made me more interested in another possibility:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;constraint failures can be discontinuous.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A model can look completely dependable on one task family and then fail dramatically when the type or interaction of requirements changes.&lt;/p&gt;

&lt;p&gt;That behavior matters for agents because developers often assume that a model that handled five previous instructions correctly will probably handle the sixth one too.&lt;/p&gt;

&lt;p&gt;ConstraintBench gives us a way to test that assumption rather than rely on it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bigger does not automatically mean more dependable
&lt;/h3&gt;

&lt;p&gt;One of the motivations for testing models from different families and capability levels was to see whether general model strength automatically translated into constraint reliability.&lt;/p&gt;

&lt;p&gt;The results suggest that this relationship is not something we should simply assume.&lt;/p&gt;

&lt;p&gt;High reasoning capability is valuable, but a production system may care about a different question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can this model reliably preserve &lt;em&gt;all&lt;/em&gt; of the requirements my application depends on?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A sophisticated answer that violates the interface contract can be less useful than a simpler answer that satisfies it perfectly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Small constraints can have large consequences
&lt;/h3&gt;

&lt;p&gt;Some benchmark requirements deliberately look mundane:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;return exactly three items;&lt;/li&gt;
&lt;li&gt;preserve this ordering;&lt;/li&gt;
&lt;li&gt;omit this value;&lt;/li&gt;
&lt;li&gt;use these exact fields;&lt;/li&gt;
&lt;li&gt;follow this conditional rule.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Individually, they can look less impressive than a reasoning problem.&lt;/p&gt;

&lt;p&gt;But in production systems, these are often the instructions that determine whether the result is usable.&lt;/p&gt;

&lt;p&gt;If an agent calculates the correct answer but places it in an invalid payload, a downstream service may reject it.&lt;/p&gt;

&lt;p&gt;If a model ignores an exclusion requirement, information that should have remained private may appear in the response.&lt;/p&gt;

&lt;p&gt;So the benchmark made me think differently about what constitutes a "minor" instruction-following failure.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Surprised Me
&lt;/h2&gt;

&lt;p&gt;The biggest surprise was how useful the &lt;strong&gt;failure pattern&lt;/strong&gt; became compared with the overall ranking.&lt;/p&gt;

&lt;p&gt;I started the project expecting to compare models primarily by their final scores.&lt;/p&gt;

&lt;p&gt;Instead, I became more interested in questions like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which task family caused a model to fail?&lt;/li&gt;
&lt;li&gt;Did it preserve high-priority requirements?&lt;/li&gt;
&lt;li&gt;Was the reasoning wrong, or was the reasoning correct but the format wrong?&lt;/li&gt;
&lt;li&gt;Did it omit something required?&lt;/li&gt;
&lt;li&gt;Did it include something prohibited?&lt;/li&gt;
&lt;li&gt;Did failure appear only when several constraints interacted?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those questions are much closer to the decisions developers make when selecting a model for a real application.&lt;/p&gt;

&lt;p&gt;The benchmark therefore changed the question I wanted to answer.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Which model is smartest?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I became more interested in:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Which model remains dependable when my application gives it many things it cannot forget?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  Why This Matters for Agents
&lt;/h2&gt;

&lt;p&gt;Constraint reliability becomes particularly important when language models stop being only conversational systems and start taking actions.&lt;/p&gt;

&lt;p&gt;An agent may need to:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;understand a request;&lt;/li&gt;
&lt;li&gt;reason about it;&lt;/li&gt;
&lt;li&gt;obey business rules;&lt;/li&gt;
&lt;li&gt;protect sensitive information;&lt;/li&gt;
&lt;li&gt;generate valid tool arguments;&lt;/li&gt;
&lt;li&gt;satisfy an API schema;&lt;/li&gt;
&lt;li&gt;preserve user requirements;&lt;/li&gt;
&lt;li&gt;decide whether an action is permitted.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A response that is "mostly correct" may not be sufficient.&lt;/p&gt;

&lt;p&gt;The system needs the model to preserve the whole contract.&lt;/p&gt;

&lt;p&gt;That is why I think instruction-following reliability deserves to be measured separately from general knowledge and reasoning ability.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I'd Measure Next
&lt;/h2&gt;

&lt;p&gt;ConstraintBench is a starting point rather than an exhaustive test.&lt;/p&gt;

&lt;p&gt;The next experiment I would run is &lt;strong&gt;explicit constraint priority&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Suppose a prompt labels its requirements:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;CRITICAL&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;HIGH&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;OPTIONAL&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the model cannot satisfy everything, does it intelligently sacrifice an optional formatting requirement before violating a privacy requirement?&lt;/p&gt;

&lt;p&gt;That would turn constraint following into a more interesting question than simply counting failures.&lt;/p&gt;

&lt;p&gt;I would also extend the benchmark to test:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;constraint density&lt;/strong&gt; — comparing 3, 5, 8, 10, and even more simultaneous requirements;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;prompt ordering&lt;/strong&gt; — whether requirements near the beginning or end of a prompt survive more reliably;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;conflicting constraints&lt;/strong&gt; — what a model sacrifices when every instruction cannot be satisfied;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;repeated constraints&lt;/strong&gt; — whether repeating critical requirements actually improves reliability;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;self-correction&lt;/strong&gt; — whether models can identify their own constraint violations;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;structured schemas vs. natural language&lt;/strong&gt; — whether explicit schemas improve compliance;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;adversarial distractions&lt;/strong&gt; — whether irrelevant instructions cause important requirements to disappear;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;latency and token cost&lt;/strong&gt; — whether additional reliability is worth the additional inference cost.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The priority experiment interests me most.&lt;/p&gt;

&lt;p&gt;If a model has to forget something, I want to know whether it forgets:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Use exactly three bullet points."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;or:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Do not expose this private value."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Those failures should not be treated as equivalent.&lt;/p&gt;




&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;The full &lt;strong&gt;ConstraintBenchmark&lt;/strong&gt; is available on Kaggle and contains the task definitions, deterministic evaluation logic, model runs, and leaderboard:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Kaggle Benchmark:&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
&lt;a href="https://www.kaggle.com/benchmarks/elshahaby/constraintbenckmark/" rel="noopener noreferrer"&gt;https://www.kaggle.com/benchmarks/elshahaby/constraintbenckmark/&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It contains &lt;strong&gt;24 deterministic cases across four task families&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Structured Extraction&lt;/li&gt;
&lt;li&gt;Reasoning Under Constraints&lt;/li&gt;
&lt;li&gt;Transformation &amp;amp; Editing&lt;/li&gt;
&lt;li&gt;Priority / Safety Preservation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The benchmark is designed to be reproducible and extensible, so additional models and new constraint categories can be evaluated using the same methodology.&lt;/p&gt;

&lt;p&gt;Building it changed the way I think about model evaluation.&lt;/p&gt;

&lt;p&gt;Raw capability matters.&lt;/p&gt;

&lt;p&gt;Reasoning matters.&lt;/p&gt;

&lt;p&gt;Knowledge matters.&lt;/p&gt;

&lt;p&gt;But when models are connected to tools, APIs, workflows, and real decisions, another property becomes just as important:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;reliability under constraints.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A model is not dependable merely because it knows the right answer.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;It also has to remember everything it was told not to get wrong.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
