DEV Community

Jeffrey Turov
Jeffrey Turov

Posted on

I benchmarked what frontier models actually know about 2026 — most of it, they don't

Kaggle Benchmarking Challenge Submission

Benchmark: https://www.kaggle.com/benchmarks/tasks/jeffreyturov/post-cutoff-150
Dataset: https://www.kaggle.com/datasets/jeffreyturov/post-cutoff-knowledge-150

The itch

Every LLM leaderboard tells you how models score on knowledge from their training data.
I wanted the opposite: what do models know about things that happened AFTER their
training cutoff?
Not retrieval-augmented answers — raw parametric knowledge of
2025-2026 facts.

The benchmark: Post-Cutoff Knowledge 150

150 factual questions built from a 7.2M-page web index, restricted to pages citing
2025-2026 events. Two construction guarantees make contamination impossible:

  1. Span-verified: each gold answer is an EXACT span of the source page's lead (deterministic string check, no LLM judge).
  2. Contamination-filtered: every candidate question is asked to a 2024-cutoff 3B model with NO context. If the small model can answer it from memory, the question is REJECTED (~4% of candidates were). Kept questions are provably outside 2024 training data: the 3B control scores 0/150.

Scoring: case-insensitive, token-boundary containment of the gold span — a lower
bound (morphological variants of a correct answer can be missed). Prompts demand
a bare fact, no sentence.

What I ran it against

9 frontier models, 150 questions each, single run, no retries on the scored attempt
(prompt: "answer with just the requested fact, no sentence"). k = correct answers
out of 150; CI = Wilson 95%.

Model Accuracy k/150 Wilson 95% CI
Gemini 3.7 Flash 24.0% 36 [17.9%, 31.4%]
DeepSeek-R1 (0528) 20.0% 30 [14.4%, 27.1%]
Gemini 2.5 Pro 16.0% 24 [11.1%, 22.4%]
GPT-5.4 11.3% 17 [7.2%, 17.4%]
Gemini 3.5 Flash-Lite 8.7% 13 [5.1%, 14.3%]
Gemini 2.5 Flash 7.3% 11 [4.1%, 12.7%]
Gemma 4 31B (open weights) 6.0% 9 [3.2%, 11.0%]
Claude Haiku 4.5 5.3% 8 [2.7%, 10.2%]
GPT-5.4 nano 3.3% 5 [1.4%, 7.6%]

Claude Sonnet 4.5 was scheduled three times and never started (infrastructure-side,
not a scoring failure) — it is excluded rather than reported as zero. Gemma 4 31B
hit per-question timeouts on 5 of 150 items (slow open-weights serving); the
unanswered items are counted as wrong, standard practice — 9/150 = 6.0%.

What surprised me

  1. The best model still fails 3 questions out of 4. Gemini 3.7 Flash leads at
    24% — meaning even the freshest frontier model has no parametric trace of most
    2025-2026 facts. Anyone building on "the model probably knows" is wrong 76% of
    the time.

  2. Reasoning does not rescue knowledge. DeepSeek-R1 (20.0%) scores second and
    beats several newer generalist models — but its long chains cannot invent a fact
    that was never in training. Reasoning moves the needle on problems, not on
    missing data.

  3. Size and price do not order the ranking. GPT-5.4 (11.3%) sits below
    DeepSeek-R1 and far below Gemini 3.7 Flash; Gemini 2.5 Pro (16.0%) beats its own
    family's newer Flash-Lite (8.7%). Freshness of training data matters more than
    benchmark muscle.

  4. The honest-control design works: the 2024-cutoff 3B control scores exactly
    0/150 — by construction — which makes every point above zero a genuine
    post-cutoff signal, not contamination.

Honest limits

  • Questions are single-hop by construction (generated from one page lead). This benchmark does not test multi-hop reasoning.
  • The "3B fails" filter makes the control ≈0 BY DESIGN — the number that matters is each frontier model's absolute score, not the gap to control.
  • Span containment is a lower bound on true accuracy.
  • 150 items → Wilson CI ≈ ±8 points. Direction is reliable, fine ranking is not.

What's next

  • Extend to 500 questions to tighten CIs.
  • Multi-hop variant: questions whose answer requires joining TWO 2025-2026 pages.
  • RAG arm: same questions WITH the source page as context, to measure the retrieval uplift per model (preliminary local data on a 3B pipeline: 0% → 33.3%).

Built with the kaggle-benchmarks library. Task source is public on the
benchmark page — fork it and run your own lineup.

Top comments (0)