This is a submission for the Kaggle Benchmarking Challenge
ConstraintBench: What Happens When an AI Has Too Many Instructions?
Most AI benchmarks ask a familiar question:
Did the model get the answer right?
I wanted to ask a slightly different one:
What happens when getting the answer right isn't enough?
Real-world prompts rarely contain one clean instruction. An AI agent may need to solve a problem, return an exact structure, include required information, exclude sensitive information, respect ordering and length requirements, perform calculations, and obey higher-priority rules—all in the same response.
A model can therefore produce an answer that looks correct while still failing the application around it.
That is what I built ConstraintBench to measure.
What I Benchmarked
ConstraintBench measures multi-constraint instruction following under specification pressure.
Instead of testing only whether a model knows the correct answer, each case combines several requirements that must be satisfied simultaneously.
The benchmark contains 24 deterministic cases across four task families, with six cases in each family.
1. Structured Extraction
These tasks require models to extract information while simultaneously satisfying requirements such as:
- exact output structure;
- mandatory fields;
- prohibited content;
- ordering;
- formatting;
- numerical conditions.
A response can contain all of the right information and still fail if it cannot be consumed reliably by another system.
2. Reasoning Under Constraints
These cases combine reasoning or numerical problems with additional output requirements.
The model has to solve the underlying problem and preserve the surrounding specification.
This lets the benchmark distinguish between:
"The model knew the answer."
and:
"The model completed the task."
Those are not always the same thing.
3. Transformation & Editing
These tasks ask models to transform content while preserving specific information and satisfying multiple editing requirements.
This resembles many real AI workflows: summarization, rewriting, document processing, data preparation, and agent-generated content.
The challenge is not simply producing fluent text. The model must change exactly what it was asked to change without accidentally changing or losing something else.
4. Priority / Safety Preservation
The final family tests whether important requirements survive when they compete with other instructions.
These cases include constraints involving privacy, prohibited information, conditional behavior, and higher-priority requirements.
This category interested me especially because not every constraint has the same consequence.
Dropping a formatting requirement is inconvenient.
Dropping a privacy or safety requirement can be much more serious.
How ConstraintBench Works
Each task family contains six cases with increasing specification pressure.
The cases combine constraints such as:
- produce valid structured output;
- include required information;
- exclude prohibited information;
- use an exact number of items;
- preserve a specified ordering;
- respect length requirements;
- perform numerical reasoning;
- follow conditional instructions;
- preserve higher-priority requirements.
ConstraintBench does not use another language model as the judge.
Instead, the cases use deterministic checks.
A required key exists or it does not.
A prohibited value appears or it does not.
A numerical condition is satisfied or it is not.
A requested structure is valid or it is not.
That makes failures easier to reproduce and interpret.
The central question behind the benchmark is:
What does a model forget first when it has too many things it cannot forget?
Models Tested
I deliberately tested models from several different families rather than comparing only closely related frontier models.
My Kaggle benchmark includes models from the Gemini, Gemma, Claude, GPT, Grok, GLM, DeepSeek, and Qwen families.
The lineup includes:
- Gemini 3.7 Flash
- Gemma 4 26B A4B
- Grok 4.20 Reasoning
- Claude Haiku 4.5
- Claude Sonnet 4.6
- Claude Opus 4.6
- GPT-5.4
- GLM-5
- DeepSeek-R1
- Qwen 3 235B A22B
I wanted this diversity because ConstraintBench is not intended to be another test of raw model intelligence.
A smaller or faster model might be extremely dependable at structured compliance.
A much larger reasoning model might solve a difficult problem correctly but miss an apparently minor output requirement.
For agents and automated workflows, those differences can matter.
Findings
The most interesting result was not simply which model appeared at the top of the leaderboard.
It was how differently models behaved across the four task families.
Correctness and compliance are different capabilities
ConstraintBench repeatedly highlights a distinction that conventional accuracy metrics can hide:
A correct answer can still be a failed response.
Imagine an agent correctly calculates a transaction amount but produces invalid JSON.
The reasoning was correct.
The workflow still breaks.
Or imagine a model produces an excellent summary but includes a piece of information it was explicitly instructed to omit.
The text may look good.
The specification was still violated.
ConstraintBench treats those requirements as part of correctness rather than as optional presentation details.
Reliability can be task-specific
Another important observation was that performance could change sharply between categories.
A model that handles priority preservation successfully does not necessarily show the same reliability on reasoning-under-constraints or transformation tasks.
Likewise, strong reasoning ability does not automatically imply perfect structured-output compliance.
This suggests that a single overall model score can hide something important.
Models have something closer to a constraint-reliability profile.
For a production application, the shape of that profile may matter more than a small difference in an aggregate benchmark score.
Models do not necessarily fail gracefully
I originally expected degradation to be relatively smooth.
A model might satisfy almost everything under light specification pressure, then gradually miss more requirements as prompts become more complicated.
The benchmark results made me more interested in another possibility:
constraint failures can be discontinuous.
A model can look completely dependable on one task family and then fail dramatically when the type or interaction of requirements changes.
That behavior matters for agents because developers often assume that a model that handled five previous instructions correctly will probably handle the sixth one too.
ConstraintBench gives us a way to test that assumption rather than rely on it.
Bigger does not automatically mean more dependable
One of the motivations for testing models from different families and capability levels was to see whether general model strength automatically translated into constraint reliability.
The results suggest that this relationship is not something we should simply assume.
High reasoning capability is valuable, but a production system may care about a different question:
Can this model reliably preserve all of the requirements my application depends on?
A sophisticated answer that violates the interface contract can be less useful than a simpler answer that satisfies it perfectly.
Small constraints can have large consequences
Some benchmark requirements deliberately look mundane:
- return exactly three items;
- preserve this ordering;
- omit this value;
- use these exact fields;
- follow this conditional rule.
Individually, they can look less impressive than a reasoning problem.
But in production systems, these are often the instructions that determine whether the result is usable.
If an agent calculates the correct answer but places it in an invalid payload, a downstream service may reject it.
If a model ignores an exclusion requirement, information that should have remained private may appear in the response.
So the benchmark made me think differently about what constitutes a "minor" instruction-following failure.
What Surprised Me
The biggest surprise was how useful the failure pattern became compared with the overall ranking.
I started the project expecting to compare models primarily by their final scores.
Instead, I became more interested in questions like:
- Which task family caused a model to fail?
- Did it preserve high-priority requirements?
- Was the reasoning wrong, or was the reasoning correct but the format wrong?
- Did it omit something required?
- Did it include something prohibited?
- Did failure appear only when several constraints interacted?
Those questions are much closer to the decisions developers make when selecting a model for a real application.
The benchmark therefore changed the question I wanted to answer.
Instead of:
Which model is smartest?
I became more interested in:
Which model remains dependable when my application gives it many things it cannot forget?
Why This Matters for Agents
Constraint reliability becomes particularly important when language models stop being only conversational systems and start taking actions.
An agent may need to:
- understand a request;
- reason about it;
- obey business rules;
- protect sensitive information;
- generate valid tool arguments;
- satisfy an API schema;
- preserve user requirements;
- decide whether an action is permitted.
A response that is "mostly correct" may not be sufficient.
The system needs the model to preserve the whole contract.
That is why I think instruction-following reliability deserves to be measured separately from general knowledge and reasoning ability.
What I'd Measure Next
ConstraintBench is a starting point rather than an exhaustive test.
The next experiment I would run is explicit constraint priority.
Suppose a prompt labels its requirements:
- CRITICAL
- HIGH
- OPTIONAL
If the model cannot satisfy everything, does it intelligently sacrifice an optional formatting requirement before violating a privacy requirement?
That would turn constraint following into a more interesting question than simply counting failures.
I would also extend the benchmark to test:
- constraint density — comparing 3, 5, 8, 10, and even more simultaneous requirements;
- prompt ordering — whether requirements near the beginning or end of a prompt survive more reliably;
- conflicting constraints — what a model sacrifices when every instruction cannot be satisfied;
- repeated constraints — whether repeating critical requirements actually improves reliability;
- self-correction — whether models can identify their own constraint violations;
- structured schemas vs. natural language — whether explicit schemas improve compliance;
- adversarial distractions — whether irrelevant instructions cause important requirements to disappear;
- latency and token cost — whether additional reliability is worth the additional inference cost.
The priority experiment interests me most.
If a model has to forget something, I want to know whether it forgets:
"Use exactly three bullet points."
or:
"Do not expose this private value."
Those failures should not be treated as equivalent.
My Benchmark
The full ConstraintBenchmark is available on Kaggle and contains the task definitions, deterministic evaluation logic, model runs, and leaderboard:
Kaggle Benchmark:
https://www.kaggle.com/benchmarks/elshahaby/constraintbenckmark/
It contains 24 deterministic cases across four task families:
- Structured Extraction
- Reasoning Under Constraints
- Transformation & Editing
- Priority / Safety Preservation
The benchmark is designed to be reproducible and extensible, so additional models and new constraint categories can be evaluated using the same methodology.
Building it changed the way I think about model evaluation.
Raw capability matters.
Reasoning matters.
Knowledge matters.
But when models are connected to tools, APIs, workflows, and real decisions, another property becomes just as important:
reliability under constraints.
A model is not dependable merely because it knows the right answer.
It also has to remember everything it was told not to get wrong.
Top comments (0)