<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: ZeroGam1ng</title>
    <description>The latest articles on DEV Community by ZeroGam1ng (@zerogam1ng).</description>
    <link>https://dev.to/zerogam1ng</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4146212%2F26e6c864-ae92-4c01-ac20-824b4b238b89.jpg</url>
      <title>DEV Community: ZeroGam1ng</title>
      <link>https://dev.to/zerogam1ng</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/zerogam1ng"/>
    <language>en</language>
    <item>
      <title>I Benchmarked Whether AI Models Forget Corrections. The Answer Surprised Me.</title>
      <dc:creator>ZeroGam1ng</dc:creator>
      <pubDate>Mon, 28 Sep 2026 13:08:00 +0000</pubDate>
      <link>https://dev.to/zerogam1ng/i-benchmarked-whether-ai-models-forget-corrections-the-answer-surprised-me-4k53</link>
      <guid>https://dev.to/zerogam1ng/i-benchmarked-whether-ai-models-forget-corrections-the-answer-surprised-me-4k53</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/kaggle-2026-09-23"&gt;Kaggle Benchmarking Challenge&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Benchmarked
&lt;/h2&gt;

&lt;p&gt;Here's the question that started this: when you correct an AI mid-conversation, does the correction stick? Or does the original wrong fact creep back in after a few thousand tokens?&lt;/p&gt;

&lt;p&gt;Psychologists call this the &lt;em&gt;continued-influence effect&lt;/em&gt; — in humans, retracted misinformation keeps influencing reasoning even after correction. I wanted to know if LLMs do the same thing, and whether it gets worse the further the correction is buried.&lt;/p&gt;

&lt;p&gt;So I built &lt;strong&gt;Temporal Memory Decay&lt;/strong&gt;: a 48-item benchmark where each item injects a synthetic fact ("the 2026 harbor gala will be in Aldermere"), explicitly corrects it ~300 tokens later ("CORRECTION: the 2026 gala will be in Vexholm"), buries both under interference, then probes which value the model reports. Scoring is graded: full credit for honoring the correction, half credit for reverting to the original, zero for anything else.&lt;/p&gt;

&lt;p&gt;Two twists make it harder than plain recall. First, the interference comes in two flavors: &lt;em&gt;unrelated&lt;/em&gt; filler versus &lt;em&gt;conflicting&lt;/em&gt; statements about the same entity with different values. Second, half the items are facts (disambiguated by year) and half are rules — "under the Meridian registry, the word 'lumen' maps to label 'jade'" — disambiguated by registry name. Delays from correction to probe run ~500, ~4,000, ~12,000, and ~32,000 tokens.&lt;/p&gt;

&lt;p&gt;Getting here took four versions. v1 was embarrassingly easy — every model scored 48/48 because the "similar" interference used different entity names, so models just pattern-matched the unique name. v3 added genuinely conflicting interference (same entity, competing values, no position hints). Flagships still aced it. v4 is the correction-decay design described above. I'm telling you about the failures because the benchmark only means something if you know what it survived.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Tested
&lt;/h2&gt;

&lt;p&gt;I ran v4 against six models, deliberately spanning the capability ladder:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;gemini-3.7-flash&lt;/strong&gt;, &lt;strong&gt;gemini-3.5-flash&lt;/strong&gt;, &lt;strong&gt;gemma-4-31b-it&lt;/strong&gt; — the flagships / strong mid-size, to establish the ceiling&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;gemini-3.5-flash-lite&lt;/strong&gt;, &lt;strong&gt;claude-haiku-4-5&lt;/strong&gt;, &lt;strong&gt;gpt-5.4-nano&lt;/strong&gt; — smaller, cheaper models, to find the breaking point&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All runs used temperature 0 (deterministic), ~600K input tokens per full run, via Kaggle's benchmark runner. Total cost across all runs: pocket change on free quota.&lt;/p&gt;

&lt;h2&gt;
  
  
  Findings
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The flagships don't forget.&lt;/strong&gt; Gemini 3.7, Gemini 3.5, and Gemma 31B each scored a perfect 48/48 — every correction honored, at every delay, under both interference types. If there's a continued-influence effect at this scale, it's invisible.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The small models cracked — but not where I was looking.&lt;/strong&gt; Here's the full scoreboard:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;th&gt;Failures&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;gemini-3.7-flash&lt;/td&gt;
&lt;td&gt;48/48 (1.000)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemini-3.5-flash&lt;/td&gt;
&lt;td&gt;48/48 (1.000)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemma-4-31b-it&lt;/td&gt;
&lt;td&gt;48/48 (1.000)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemini-3.5-flash-lite&lt;/td&gt;
&lt;td&gt;47/48 (0.979)&lt;/td&gt;
&lt;td&gt;1 reversion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;claude-haiku-4-5&lt;/td&gt;
&lt;td&gt;47/48 (0.979)&lt;/td&gt;
&lt;td&gt;1 reversion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gpt-5.4-nano&lt;/td&gt;
&lt;td&gt;38/48 (0.792)&lt;/td&gt;
&lt;td&gt;6 reversions, 4 distractor grabs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Now the interesting part: &lt;em&gt;where&lt;/em&gt; the failures happened. The nano model's 10 failures were &lt;strong&gt;all&lt;/strong&gt; in conflicting-interference rule cells — 0/3, 1/3, 1/3, 0/3 across the four delays. Every fact item: perfect. Every unrelated-interference item: perfect. And the lite and Haiku models each reverted exactly once — both in the &lt;strong&gt;identical&lt;/strong&gt; cell: 32,000-token delay × conflicting interference × rule item. Two models from two different labs, cracking at the same coordinates.&lt;/p&gt;

&lt;p&gt;Read that again: the nano model failed conflicting-rule items at &lt;strong&gt;500 tokens&lt;/strong&gt; — the shortest delay, my control condition. There is no decay curve here. Failures are flat across delays. Distance isn't the variable. &lt;em&gt;Competition is.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The insight I didn't expect:&lt;/strong&gt; it's not that corrections fade with time. It's that certain kinds of bindings shatter under competing mappings. Year-keyed facts ("the &lt;em&gt;2026&lt;/em&gt; gala") are rock solid — years are distinctive, ordered keys the model handles effortlessly. Registry-keyed rules ("under the &lt;em&gt;Meridian&lt;/em&gt; registry") fall apart when other registries map the same word to different labels. The small models can't reliably bind (word, registry) → label when distractors pile on — they either revert to the original value or grab a competitor's label. The flagships do it flawlessly, so this is a capability threshold, not a universal flaw.&lt;/p&gt;

&lt;p&gt;I went in hunting for memory decay over distance. The models told me I was asking the wrong question. The failure mode isn't "old information fades" — it's "similar information collides." That reframes how I'd think about long-context reliability: the risk isn't the &lt;em&gt;length&lt;/em&gt; of the conversation, it's the &lt;em&gt;density of competing claims&lt;/em&gt; about the same entities.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I'd measure next:&lt;/strong&gt; the obvious follow-up is a dose-response — vary the &lt;em&gt;number&lt;/em&gt; of competing registry mappings (1 vs 3 vs 7 distractors) and map the failure curve on small models. If P(failure) scales with competitor count, that's a clean, actionable law for anyone stuffing conflicting sources into a long context window. I'd also test whether distinctive keys (years, IDs) are universally robust or whether there's a competition level that breaks those too.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Honest limitations:&lt;/strong&gt; 48 items is small; these are point measurements at temperature 0, not distributions. My "facts" and "rules" differ in more ways than just the key type (sentence structure, distractor style), so the fact/rule dissociation is suggestive, not airtight. And this is a synthetic task — real conversations have messier corrections. Take it as a lens, not a verdict.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Benchmark
&lt;/h2&gt;

&lt;p&gt;🔗 &lt;strong&gt;&lt;a href="https://www.kaggle.com/benchmarks/tasks/zerogam1ng/temporal-memory-decay/8" rel="noopener noreferrer"&gt;temporal-memory-decay on Kaggle Benchmarks&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The task is public — all 48 items, the scoring code, and every model run with full per-cell breakdowns. If you want to run your favorite model against it, or remix the design (the dose-response experiment above is wide open), go for it. That's the point.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>kagglechallenge</category>
      <category>ai</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
