DEV Community

Cover image for What Does AI Forget First?
Riya Maurya
Riya Maurya

Posted on AI-assisted

What Does AI Forget First?

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

I wanted to know one thing: when a coding chat gets crowded, does an AI model keep following the rules you set earlier?

People often say that AI “forgets” things in long chats. I wanted to measure the behavior instead of just guessing. I’m deliberately not calling it memory loss, because I can’t see inside the model. What I can measure is constraint adherence under context load: does the final code still respect a rule stated earlier?

I called the experiment PriorityDecay.

  1. I set up a fictional student internship platform and stated five project constraints:
    • Backend must use Python FastAPI.
    • Database must be PostgreSQL.
    • Resume uploads must be capped at 5 MB.
    • Never hardcode or expose API keys or secrets.
    • JavaScript variables must use snake_case. I chose this intentionally because it is unusual in JavaScript and can help distinguish explicit instruction-following from the model’s default style.
  2. I added 0, 10, or 20 unrelated side topics, such as Git questions, README typos, REST vs. GraphQL, and stand-up tips.
  3. I then asked for a resume-upload endpoint and a small upload widget without repeating the original constraints.
  4. I evaluated the generated response with plain Python checks rather than an LLM judge. The checks used regular expressions and string matching, so they are useful signals—not proof that the generated code works.

Each condition—SHORT (0 distractors), MEDIUM (10), and LONG (20)—was run five times, for 15 runs total. The scorer used eight checks: the five stated constraints, two UI requirements (show the selected filename and upload status; disable the button during upload), and an additional hygiene check for credentials embedded in a default database URL.

I kept this small enough to inspect the outputs by hand.

Model Tested

I ran the benchmark on google/gemini-3.7-flash, as shown in the Kaggle notebook output.

This is a one-model pilot with 15 runs. The findings below describe this experiment, not coding models in general.

Findings

Headline numbers (score out of 8, plus retention of the five stated constraints across five runs per condition):

Condition Distractors Mean score / 8 Stated constraints passed
SHORT 0 7.0 25/25 (100%)
MEDIUM 10 6.8 22/25 (88%)
LONG 20 6.6 23/25 (92%)

Per-constraint passes out of five runs:

Constraint SHORT MEDIUM LONG
FastAPI 5 5 5
5 MB limit 5 5 5
No hardcoded API key 5 5 5
JavaScript snake_case 5 5 5
PostgreSQL evidence detected 5 2 3

The aggregate scores dipped slightly as context increased, but most of the change in the five stated constraints came from the PostgreSQL check. FastAPI, the upload limit, the explicit API-key check, and snake_case passed in all 15 runs.

What surprised me

1. The apparent “PostgreSQL decay” was often a measurement problem.

I inspected the outputs where the PostgreSQL check failed. None switched explicitly to SQLite, MySQL, or another database. In four of the five failures, the response used DATABASE_URL without explicitly naming PostgreSQL. That can be a valid configuration pattern, but the string-based check could not verify which database the environment variable would point to. So a failed check did not necessarily mean the model violated the requirement.

The reverse also happened: one LONG output passed because a comment mentioned “PostgreSQL,” and another passed because a default URL contained the term even though the endpoint did not actually use a database.

This is why a passing regex check should not be confused with verified behavior.

2. Two LONG outputs did not persist anything to a database.

In two of the five LONG runs, the endpoint validated the uploaded file and returned a response without writing a record to a database. All SHORT and MEDIUM outputs stored a record.

The final task did not explicitly require storing resume metadata, so I would not count this as a proven violation of the original requirements. Still, it is a useful implementation difference to investigate in a follow-up benchmark where persistence is stated as an explicit acceptance criterion.

3. The credential check told a different story from the API-key check.

The extra check flagged default database URLs with embedded credentials in 5/5 SHORT outputs, 2/5 MEDIUM outputs, and 2/5 LONG outputs. In other words, the longer-context outputs more often used environment-based configuration rather than credentials embedded in the URL.

At the same time, the explicit “don’t hardcode API keys” check passed in all 15 runs, while the database-URL check flagged embedded credentials in 9/15. These are different checks for different kinds of secrets; passing one does not establish that a solution is secure. With five runs per condition, I’m not claiming the difference is a reliable effect.

4. My scorer was wrong sometimes—and that matters.

The filename-and-status UI check failed in one MEDIUM output and three LONG outputs. I manually inspected all four and found that each visibly displayed the selected filename and a status message. The outputs used element IDs that my check did not expect.

So the apparent UI decline was a measurement artifact, not a demonstrated model failure. If I had only looked at the leaderboard score, I would have told the wrong story.

5. I didn’t find convincing evidence of “priority decay.”

At this scale, I found no clear evidence that the stated constraints reliably deteriorated as distractors increased. Four of the five constraints passed in every run, and the PostgreSQL check had an important observability limitation.

That is still a useful result: it shows how easy it is to mistake limitations in a benchmark’s evaluator for limitations in the model.

What I’d Measure Next

  • A no-constraint control for every task. Without it, I can’t separate “the model retained the rule” from “the model would have done this anyway.” An earlier informal run on a previous version of the task produced Flask, SQLite, camelCase JavaScript, and no upload-size limit. That is suggestive, but it was only one run and is not a controlled baseline for these results.
  • Much longer contexts. Twenty bundled distractor topics are still a short chat compared with a real development session.
  • Explicit persistence requirements. If database storage matters, the task should say exactly what must be persisted and how it will be tested.
  • Conflicting later instructions. For example, an earlier “never log auth tokens” followed by a request for verbose request logging could test whether a model preserves important security rules.
  • Constraint position and phrasing. I did not vary whether a rule appeared at the beginning, middle, or end, or whether it was phrased positively or negatively.
  • A stronger evaluator. Parse code with an AST, run the endpoint in a test harness, and add human review for ambiguous cases. I’d also test more models and use more than five runs per condition.

Limitations, Plainly

This was one model, one task, 15 runs, a short maximum context, and string-based checks that missed some valid behavior. There was no controlled no-constraint baseline in the reported results.

The honest conclusion is: this pilot found only a small change in aggregate scores, and some of the apparent failures came from the evaluator itself. It does not show that models never forget, nor does it establish that longer context causes degradation.

My Benchmark

Kaggle notebook: [https://www.kaggle.com/code/riyaax07/new-benchmark-task-30fbf]

The experiment and results analysis are available in the notebook. I’d be interested in seeing whether the pattern changes with other models, longer contexts, and execution-based tests.

The question I want to investigate next is simple:

When the context gets longer, does a coding model still follow the requirements that make your project your project?

Top comments (1)

Collapse
 
ahmetozel profile image
Ahmet Özel •

The DATABASE_URL cases show why a binary regex verdict is too strong for this task. There are at least three outcomes: PostgreSQL use demonstrated, another database used, and configuration left unresolved. Treating the last as a violation or as a pass both overstate the evidence.

For the execution-based follow-up, I would provide the same PostgreSQL connection in every condition and inspect the actual connection and stored record. Keep evaluator uncertainty separate from implementation failure in the pilot table. That preserves the useful negative result without letting a comment containing PostgreSQL count as database behavior.