This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
In drug-discovery target evaluation, the hardest discipline is not letting hindsight leak into a judgment that's supposed to be made at an earlier point in time. Given a piece of clinical/genetic evidence and a decision date, the only question that should matter is: was this evidence actually available before that date? Everything else, including whether we now know the drug succeeded or failed, is irrelevant, and if it sneaks in, it invalidates the judgment.
That's a simple rule (evidence_date < decision_date), but I wanted to test whether LLMs apply it mechanically even when the evidence content is emotionally/narratively loaded, i.e. describes a drug whose eventual real-world fate the model may already "know" from training data.
I built a 17-item eligibility-classification benchmark from a firsthand-verified evidence graph covering two diseases (Crohn's disease, axial spondyloarthritis) and three drug classes (anti-IL-17A, anti-IL-23p19, anti-IL-12/23p40): the same three molecules, opposite clinical outcomes in each disease, cross-checked against ClinicalTrials.gov and PubMed.
- 10 plain WITHIN_CUTOFF items: evidence clearly predates the decision.
- 2 hindsight-trap items: evidence postdates the cutoff and is thematically reinforcing (not a dramatic "trial succeeded/failed" headline, but a mechanistic paper whose conclusion agrees with the claim's own direction, the subtler, more realistic trap).
- 5 boundary items: evidence dated the exact same day as the cutoff. The rule is strict-less-than, so same-day does not count as available; this tests whether models default to "same day is close enough."
Models Tested
Five models chosen for contrast, not just coverage:
-
google/gemini-2.5-flashandgoogle/gemini-2.5-pro: a smaller/larger pair from the same family, to see whether scale alone would matter for this failure mode. -
openai/gpt-oss-20b: the smallest, open-weight model available to me, picked as the one most likely to slip. -
anthropic/claude-haiku-4-5: a small model from a different vendor family, to separate "small models slip" from "this particular small model slips." -
deepseek-ai/deepseek-r1-0528: a reasoning model, to test the opposite worry: does more deliberate step-by-step reasoning actually make associative/hindsight leakage more likely, not less?
I also ran every model twice: once with an explicit instruction not to use outcome knowledge, and once with that instruction removed, to isolate how much of any good result was the instruction doing the work versus the model's own judgment.
Findings
Three of five models were flawless, every single time. gemini-2.5-flash, gemini-2.5-pro, and claude-haiku-4-5 scored a perfect 17/17 on every sub-metric, with and without the explicit guardrail instruction, across three independent re-runs of the whole benchmark. Not one wobble.
The smallest open-weight model was not flawless, and not even consistent with itself.
Three independent full runs of gpt-oss-20b against the identical prompts produced three different scorecards:
| run | with-guardrail overall | trap (2 items) | boundary (5 items) | no-guardrail overall | trap | boundary |
|---|---|---|---|---|---|---|
| 1 | 94.1% | 50% | 100% | 94.1% | 50% | 100% |
| 2 | 88.2% | 100% | 60% | 82.4% | 0% | 80% |
| 3 | 76.5% | 50% | 40% | 82.4% | 50% | 60% |
| mean | 86.3% | 66.7% | 66.7% | 86.3% | 33.3% | 80% |
This is the finding I didn't expect going in: the failure mode isn't a stable "this model falls for hindsight" pattern, it's instability itself. The same model, the same prompt, answering the same question about the same evidence, gave a different verdict depending on nothing I controlled. For a task where the entire point is "give the same principled answer every time regardless of how tempting the content is," that unpredictability is arguably a worse failure than a consistent bias would be: you can correct for a consistent bias; you can't correct for a coin flip.
The reasoning model didn't fail the judgment, it failed the pipeline.
deepseek-r1-0528 never produced a scorable answer. Looking at the raw output, its actual reasoning was correct every time I inspected it, e.g. for one item: "the year 2006 is less than 2014 ... Therefore, the evidence was available", right answer, right reasoning. The failure was that R1 emits its full chain-of-thought inside <think>...</think> tags before the answer, and the structured-output parser tried to parse the entire response (thinking block included) as JSON and choked. This is a genuinely useful, unglamorous insight for anyone benchmarking reasoning models with structured-output tooling: you need to strip the <think> block before parsing, or the model can be right and still register as a total failure.
What surprised me overall: I built this benchmark to catch models getting seduced by dramatic content into misjudging dates. That specific failure mode never showed up, not once, in any model, in any run. What showed up instead were two failure modes I hadn't designed for: run-to-run instability in a small model, and an infrastructure mismatch with a reasoning model's output format. Both are more useful findings than the one I went looking for, because they're the kind of thing you only find by actually running the benchmark against a diverse model set instead of stopping at the first clean result.
What I'd measure next: whether gpt-oss-20b's instability is temperature-driven (rerun at temperature 0 to see if it stabilizes), and whether stripping <think> tags before parsing resolves deepseek-r1 entirely or whether its judgment quality on the trap items specifically is also worse once it's actually scorable.

My Benchmark
https://www.kaggle.com/code/jeonghosong/new-benchmark-task-fe5fd
Top comments (0)