DEV Community

Cover image for I Benchmarked Whether AI Models Forget Corrections. The Answer Surprised Me.
ZeroGam1ng
ZeroGam1ng

Posted on

I Benchmarked Whether AI Models Forget Corrections. The Answer Surprised Me.

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

Here's the question that started this: when you correct an AI mid-conversation, does the correction stick? Or does the original wrong fact creep back in after a few thousand tokens?

Psychologists call this the continued-influence effect — in humans, retracted misinformation keeps influencing reasoning even after correction. I wanted to know if LLMs do the same thing, and whether it gets worse the further the correction is buried.

So I built Temporal Memory Decay: a 48-item benchmark where each item injects a synthetic fact ("the 2026 harbor gala will be in Aldermere"), explicitly corrects it ~300 tokens later ("CORRECTION: the 2026 gala will be in Vexholm"), buries both under interference, then probes which value the model reports. Scoring is graded: full credit for honoring the correction, half credit for reverting to the original, zero for anything else.

Two twists make it harder than plain recall. First, the interference comes in two flavors: unrelated filler versus conflicting statements about the same entity with different values. Second, half the items are facts (disambiguated by year) and half are rules — "under the Meridian registry, the word 'lumen' maps to label 'jade'" — disambiguated by registry name. Delays from correction to probe run ~500, ~4,000, ~12,000, and ~32,000 tokens.

Getting here took four versions. v1 was embarrassingly easy — every model scored 48/48 because the "similar" interference used different entity names, so models just pattern-matched the unique name. v3 added genuinely conflicting interference (same entity, competing values, no position hints). Flagships still aced it. v4 is the correction-decay design described above. I'm telling you about the failures because the benchmark only means something if you know what it survived.

Models Tested

I ran v4 against six models, deliberately spanning the capability ladder:

  • gemini-3.7-flash, gemini-3.5-flash, gemma-4-31b-it — the flagships / strong mid-size, to establish the ceiling
  • gemini-3.5-flash-lite, claude-haiku-4-5, gpt-5.4-nano — smaller, cheaper models, to find the breaking point

All runs used temperature 0 (deterministic), ~600K input tokens per full run, via Kaggle's benchmark runner. Total cost across all runs: pocket change on free quota.

Findings

The flagships don't forget. Gemini 3.7, Gemini 3.5, and Gemma 31B each scored a perfect 48/48 — every correction honored, at every delay, under both interference types. If there's a continued-influence effect at this scale, it's invisible.

The small models cracked — but not where I was looking. Here's the full scoreboard:

Model Score Failures
gemini-3.7-flash 48/48 (1.000) —
gemini-3.5-flash 48/48 (1.000) —
gemma-4-31b-it 48/48 (1.000) —
gemini-3.5-flash-lite 47/48 (0.979) 1 reversion
claude-haiku-4-5 47/48 (0.979) 1 reversion
gpt-5.4-nano 38/48 (0.792) 6 reversions, 4 distractor grabs

Now the interesting part: where the failures happened. The nano model's 10 failures were all in conflicting-interference rule cells — 0/3, 1/3, 1/3, 0/3 across the four delays. Every fact item: perfect. Every unrelated-interference item: perfect. And the lite and Haiku models each reverted exactly once — both in the identical cell: 32,000-token delay × conflicting interference × rule item. Two models from two different labs, cracking at the same coordinates.

Read that again: the nano model failed conflicting-rule items at 500 tokens — the shortest delay, my control condition. There is no decay curve here. Failures are flat across delays. Distance isn't the variable. Competition is.

The insight I didn't expect: it's not that corrections fade with time. It's that certain kinds of bindings shatter under competing mappings. Year-keyed facts ("the 2026 gala") are rock solid — years are distinctive, ordered keys the model handles effortlessly. Registry-keyed rules ("under the Meridian registry") fall apart when other registries map the same word to different labels. The small models can't reliably bind (word, registry) → label when distractors pile on — they either revert to the original value or grab a competitor's label. The flagships do it flawlessly, so this is a capability threshold, not a universal flaw.

I went in hunting for memory decay over distance. The models told me I was asking the wrong question. The failure mode isn't "old information fades" — it's "similar information collides." That reframes how I'd think about long-context reliability: the risk isn't the length of the conversation, it's the density of competing claims about the same entities.

What I'd measure next: the obvious follow-up is a dose-response — vary the number of competing registry mappings (1 vs 3 vs 7 distractors) and map the failure curve on small models. If P(failure) scales with competitor count, that's a clean, actionable law for anyone stuffing conflicting sources into a long context window. I'd also test whether distinctive keys (years, IDs) are universally robust or whether there's a competition level that breaks those too.

Honest limitations: 48 items is small; these are point measurements at temperature 0, not distributions. My "facts" and "rules" differ in more ways than just the key type (sentence structure, distractor style), so the fact/rule dissociation is suggestive, not airtight. And this is a synthetic task — real conversations have messier corrections. Take it as a lens, not a verdict.

My Benchmark

🔗 temporal-memory-decay on Kaggle Benchmarks

The task is public — all 48 items, the scoring code, and every model run with full per-cell breakdowns. If you want to run your favorite model against it, or remix the design (the dose-response experiment above is wide open), go for it. That's the point.

Top comments (0)