This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
Assistants remember things about us now. The failure I kept noticing is not forgetting. It's remembering the old version. You mention you moved, switched jobs or went vegan, and a few weeks later the assistant happily plans around the person you used to be.
So I built Stale Facts: 34 conversation histories where one fact about the user changes partway through, followed by questions about it. There are five kinds of question:
| Type | What it asks | Example |
|---|---|---|
| CURRENT | what is true now | "Which city am I living in these days?" |
| HISTORICAL | what was true at a past date | "Where was I living in January 2026?" |
| PRESUPPOSED | a request that quietly assumes the old fact | "Any cafes with good wifi near my flat in Hyderabad?" |
| ABSTAIN | the change was only a rumour | "What's my rent right now? I'm filling in my HRA form." |
| CONTROL | something nearby never changed, or a planned change was called off | "Which city's office do I work out of?" |
The whole history sits in the model's context window. That was on purpose. Memory products usually fail at retrieval, so I wanted to know what happens when retrieval is perfect. If a model gets it wrong with the answer right there in the transcript, a better retriever won't fix it.
The histories are 5 to 8 dated conversations, most of them about something else entirely (a leaky tap, a Coorg trip, a thesis intro). The changing fact is rarely the topic. On top of that:
- the old value often comes back after the change: a trip home, a refund from the old ISP, lunch with ex-colleagues
- some changes are only implied, never announced ("the gemeente appointment for my BSN is on the 26th", nothing about moving to Amsterdam)
- some changes are announced before they take effect, so "what was my job title in December?" has the old answer even though the promotion was already mentioned
- decoys everywhere: a sister's diet, a neighbour's vet, the sales team's offsite city
I also scored two rules that use no model at all: "answer with the most recently mentioned value" and "answer with the first one mentioned". On the three example histories I started from, the first rule got every current-value question right, which told me those examples measured nothing. On the final 34 it gets 6/20, and the first-mention rule gets 8/20.
Grading is exact matching first: if the reply contains only the right value, it passes, and only the stale value, it fails. Anything less clear (both values named, a hedge, every PRESUPPOSED and ABSTAIN reply) goes to three judge models from three different labs, none of them in the lineup, and the majority wins.
Models Tested
Eleven models from seven labs, picked so each family has a big and a small model where Kaggle offers one. That turns "does a bigger model fix this?" into something I can check instead of assume.
| Lab | Models |
|---|---|
| Anthropic | Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5 |
| OpenAI | GPT-6 Astra, GPT-5.4 mini |
| Gemini 3.8 Flash, Gemma 4 31B | |
| xAI | Grok 4.20 (reasoning) |
| DeepSeek | DeepSeek-R1 |
| Alibaba | Qwen3 235B Instruct |
| Zhipu | GLM-5 |
Gemini 3.7 Flash also shows up in the results as a bonus: Kaggle runs its default model every time a task is pushed.
Judges: Gemini 2.5 Pro, GPT-5.4 and Claude Opus 4.5. They come from the same three big labs as most of the lineup, but none of them is a model under test.
Two models I wanted didn't make it. Grok 4.6 is in Kaggle's model list, but the proxy returns "model not found" for it. gpt-oss-120b kept cutting its replies off mid-word (literally "You're currently living in **") and hitting rate limits, so GLM-5 took its slot.
Findings
| Model | Now | Past date | Stale premise | Rumour | Unchanged |
|---|---|---|---|---|---|
| Gemini 3.7 Flash | 20/20 | 21/21 | 20/20 | 10/10 | 13/13 |
| Gemini 3.8 Flash | 20/20 | 21/21 | 20/20 | 10/10 | 13/13 |
| Claude Opus 5 | 20/20 | 21/21 | 20/20 | 10/10 | 13/13 |
| GPT-6 Astra | 20/20 | 21/21 | 18/20 | 10/10 | 13/13 |
| Gemma 4 31B | 20/20 | 21/21 | 17/20 | 9/9 | 13/13 |
| Claude Sonnet 5 | 19/20 | 21/21 | 17/20 | 10/10 | 13/13 |
| GLM-5 | 19/20 | 21/21 | 15/20 | 8/10 | 13/13 |
| Grok 4.20 | 20/20 | 21/21 | 7/20 | 10/10 | 13/13 |
| DeepSeek-R1 | 18/20 | 18/21 | 8/20 | 10/10 | 13/13 |
| Qwen3 235B | 19/20 | 18/21 | 8/20 | 9/10 | 13/13 |
| Claude Haiku 4.5 | 20/20 | 13/21 | 11/20 | 9/10 | 13/13 |
| GPT-5.4 mini | 18/20 | 18/21 | 0/20 | 5/10 | 13/13 |
1. Knowing the fact and acting on it are two different skills. Every model answered "what's true now?" correctly at least 90% of the time. On requests built on the old fact, the same models ranged from 0% to 100%. GPT-5.4 mini is the extreme case: 18 of 20 on the direct questions, 0 of 20 on the stale premises. It knows the user left Kestrel. Asked to "draft a short Slack message to my manager at Kestrel", it wrote "Hi Deepa", the old manager, by name.
2. It breaks on writing tasks. The two hardest items in the set were both "draft me a message" requests: a leave request to a manager who was replaced in February (3 of 12 models noticed) and an out-of-office note for a squad the user had moved off (4 of 12). A membership freeze email to the climbing gym the user left was close behind. Most models just wrote the message. Once the request becomes a writing job, the premise stops being a claim to check and turns into a blank to fill in. This is Claude Haiku 4.5, writing to a manager who was replaced back in February:
"Hi Marcus, need Oct 8-9 off for a cousin's wedding in Jaipur. Happy to wrap anything urgent before then. Thanks."
3. Implied changes are the ones that slip. Pooled across models, stale-premise accuracy was 56% when the change was only implied, against 76% when it was stated outright and 75% when it was a correction. On direct questions the same split barely moves. Models can work out an implied change when you ask about it. They just don't volunteer it.
4. Some models state the date that proves them wrong. Claude Haiku 4.5 answered 7 of 21 past-date questions with today's value, and a few of those answers contradict themselves in one breath:
Q: Which company was I working for in February 2026?
A: "Orbitra Health. You started there on March 9, 2026, after your last day at Kestrel Analytics was February 27."Q: Which gym was I climbing at regularly in December 2025?
A: "Boulderhaus. You got a membership there on January 5th, 2026, but before that you were at Vertigo Walls in Indiranagar three evenings a week."
The model has every piece of the answer and still picks the latest value.
5. Size helps inside a family, but it isn't the whole story. Within each lab the bigger model caught more stale premises: Claude Opus 5 got 20 of 20 against Haiku 4.5's 11, GPT-6 Astra 18 against GPT-5.4 mini's 0, Gemini 3.8 Flash 20 against Gemma 4's 17. Across labs the ordering falls apart. Gemini 3.8 Flash, a cheap tier, was perfect. Grok 4.20, running in reasoning mode, caught 7.
What surprised me. Grok 4.20 in reasoning mode got all 64 of the direct questions right: what's true now, what was true at a past date, which changes were only rumours, which facts never moved. Then it went along with 13 of the 20 stale premises, including "Hi Deepa" for the old manager and a week of paneer and curd lunches for a user who went vegan in March. It knows. It just doesn't check. Thinking longer doesn't help if the model never thinks to question the request.
The other end of the table is just as telling. Claude Opus 5 and both Gemini Flash models were perfect on all 84 probes, and the Gemini models flag the premise in a single friendly line before doing the task ("Just a quick heads-up: double-check if you need to send this to Chiara instead"). The fix isn't refusing or lecturing. It's one sentence, which makes the models that skip it harder to excuse.
How reliable is the grading? The deterministic matcher settled 28% of answers. For the rest, the first two judges agreed 98% of the time, so the tiebreak judge was rarely needed. A random sample of 40 judge verdicts was checked against the rubric and all 40 held up (one lenient but defensible), and every one of Grok's stale-premise verdicts was re-read because that result is the most surprising.
What I'd measure next.
- The same histories through real memory systems (a vector store, a summarizing memory) instead of the full transcript. That was the original question behind this, and the stale-premise result suggests even perfect retrieval won't be enough on its own.
- Whether one line in the system prompt ("if a request relies on something that changed in our history, say so first") closes the gap. If it does, this is a default-behaviour problem, not a capability one.
- Longer histories and more than one changing fact per history.
Limitations. 34 histories and 84 probes is small: differences of a few probes between two models are noise, and the Wilson intervals on the leaderboard say so. The histories were drafted with LLM help from a detailed spec and then reviewed one by one. One broken item was caught and fixed along the way (a past-date question about a month the history gave no evidence for).
Two Kaggle gotchas worth knowing
-
The default judge is the model being tested. Inside a Kaggle run,
kbench.judge_llmresolves to the same model askbench.llm, so every model grades its own answers unless you pin a judge. I pinned three. -
Expensive models fail with a quota error that isn't about your quota. The proxy reserves the worst-case cost of a call from the maximum output length before running it, and with the default limit a single GPT-6 Astra call reserves more than a day's allowance. Passing
extra_api_params={"max_tokens": 8192}tollm.prompt()fixes it. (max_output_tokensis rejected;max_tokensandmax_completion_tokensboth work.)
My Benchmark
- Benchmark and leaderboard: Stale Facts on Kaggle
- Tasks (one per question type): stale-current, stale-historical, stale-presupposed, stale-abstain, stale-control
- All 34 histories, the schema and the grading code: harshsingh1708/stale-facts-bench
Top comments (1)