Benchmark: https://www.kaggle.com/benchmarks/tasks/jeffreyturov/post-cutoff-150
Dataset: https://www.kaggle.com/datasets/jeffreyturov/post-cutoff-knowledge-150
The itch
Every LLM leaderboard tells you how models score on knowledge from their training data.
I wanted the opposite: what do models know about things that happened AFTER their
training cutoff? Not retrieval-augmented answers — raw parametric knowledge of
2025-2026 facts.
The benchmark: Post-Cutoff Knowledge 150
150 factual questions built from a 7.2M-page web index, restricted to pages citing
2025-2026 events. Two construction guarantees make contamination impossible:
- Span-verified: each gold answer is an EXACT span of the source page's lead (deterministic string check, no LLM judge).
- Contamination-filtered: every candidate question is asked to a 2024-cutoff 3B model with NO context. If the small model can answer it from memory, the question is REJECTED (~4% of candidates were). Kept questions are provably outside 2024 training data: the 3B control scores 0/150.
Scoring: case-insensitive, token-boundary containment of the gold span — a lower
bound (morphological variants of a correct answer can be missed). Prompts demand
a bare fact, no sentence.
What I ran it against
9 frontier models, 150 questions each, single run, no retries on the scored attempt
(prompt: "answer with just the requested fact, no sentence"). k = correct answers
out of 150; CI = Wilson 95%.
| Model | Accuracy | k/150 | Wilson 95% CI |
|---|---|---|---|
| Gemini 3.7 Flash | 24.0% | 36 | [17.9%, 31.4%] |
| DeepSeek-R1 (0528) | 20.0% | 30 | [14.4%, 27.1%] |
| Gemini 2.5 Pro | 16.0% | 24 | [11.1%, 22.4%] |
| GPT-5.4 | 11.3% | 17 | [7.2%, 17.4%] |
| Gemini 3.5 Flash-Lite | 8.7% | 13 | [5.1%, 14.3%] |
| Gemini 2.5 Flash | 7.3% | 11 | [4.1%, 12.7%] |
| Gemma 4 31B (open weights) | 6.0% | 9 | [3.2%, 11.0%] |
| Claude Haiku 4.5 | 5.3% | 8 | [2.7%, 10.2%] |
| GPT-5.4 nano | 3.3% | 5 | [1.4%, 7.6%] |
Claude Sonnet 4.5 was scheduled three times and never started (infrastructure-side,
not a scoring failure) — it is excluded rather than reported as zero. Gemma 4 31B
hit per-question timeouts on 5 of 150 items (slow open-weights serving); the
unanswered items are counted as wrong, standard practice — 9/150 = 6.0%.
What surprised me
The best model still fails 3 questions out of 4. Gemini 3.7 Flash leads at
24% — meaning even the freshest frontier model has no parametric trace of most
2025-2026 facts. Anyone building on "the model probably knows" is wrong 76% of
the time.Reasoning does not rescue knowledge. DeepSeek-R1 (20.0%) scores second and
beats several newer generalist models — but its long chains cannot invent a fact
that was never in training. Reasoning moves the needle on problems, not on
missing data.Size and price do not order the ranking. GPT-5.4 (11.3%) sits below
DeepSeek-R1 and far below Gemini 3.7 Flash; Gemini 2.5 Pro (16.0%) beats its own
family's newer Flash-Lite (8.7%). Freshness of training data matters more than
benchmark muscle.The honest-control design works: the 2024-cutoff 3B control scores exactly
0/150 — by construction — which makes every point above zero a genuine
post-cutoff signal, not contamination.
Honest limits
- Questions are single-hop by construction (generated from one page lead). This benchmark does not test multi-hop reasoning.
- The "3B fails" filter makes the control ≈0 BY DESIGN — the number that matters is each frontier model's absolute score, not the gap to control.
- Span containment is a lower bound on true accuracy.
- 150 items → Wilson CI ≈ ±8 points. Direction is reliable, fine ranking is not.
What's next
- Extend to 500 questions to tighten CIs.
- Multi-hop variant: questions whose answer requires joining TWO 2025-2026 pages.
- RAG arm: same questions WITH the source page as context, to measure the retrieval uplift per model (preliminary local data on a 3B pipeline: 0% → 33.3%).
Built with the kaggle-benchmarks library. Task source is public on the
benchmark page — fork it and run your own lineup.
Top comments (0)