This is a submission for the Kaggle Benchmarking Challenge.
🎥 3-minute overview (narrated):
What I Benchmarked
Standards have a long tail of proposals that never shipped: emoji withdrawn from a draft
list, PEPs that were rejected or deferred, IETF Internet-Drafts that expired, HTTP headers
that never made it into the registry. They are all over the web: blog posts, mailing lists,
early implementations. So a model has seen plenty of text about them, often written as if
they were real.
I wanted to know: do LLMs treat these dead proposals as real standards? And, more
importantly, does it show up in the things models produce, not only when you ask them
directly? A model that says "PEP 505 was deferred" but then writes a?.b in your code is
the dangerous case.
Drafts That Never Shipped asks about 122 entities in five families:
| Family | Never shipped (example) | Matched shipped "twin" |
|---|---|---|
| Emoji | Apple Core (withdrawn from the Emoji 17.0 draft) | Lime, Shaking Face |
| PEPs | PEP 505 None-aware ?. (Deferred), PEP 3150 given (Deferred) |
PEP 572 :=, PEP 634 match
|
| PEP draft syntax | PEP 622's case .RED:, PEP 750's first-draft greet"..."
|
case Color.RED:, t"..."
|
| IETF drafts | RateLimit headers draft, JSON Schema, OAuth 2.1 | RFC 8941, RFC 9068, … |
| HTTP fields |
Idempotency-Key, RateLimit-Policy, Want-Digest (obsoleted) |
Priority, Cache-Status, Prefer
|
Plus fictional placebos (fake emoji names, PEP numbers that don't exist, draft names that
404 on Datatracker) to measure plain hallucination.
Each entity is asked about on several surfaces:
-
Status: "Did X ship?" →
{"shipped": true|false} -
Lookup: the codepoint / Python version / RFC number, or
NONE -
Artefact: actually use it:
- a chat message with the emoji;
- a Python 3.14 snippet;
- a citation;
- raw HTTP headers.
Every artefact prompt offers a neutral opt-out (NOT_AVAILABLE), so a correct answer
always exists.
Ground truth is never hand-typed. Every label comes from a pinned, hashed,
machine-readable source:
- Unicode
emoji-test.txt/UnicodeData.txt(15.1 → 18.0); -
peps.json; - the IETF Datatracker API;
- the IANA HTTP Field Name Registry.
Scoring is fully deterministic, with no LLM judge. Python snippets are compiled (never
executed) under Python 3.14: if your "PEP 505 example" doesn't compile, that's
evidence you believed the syntax exists. My own curation was wrong once (I assumed the
HTTP QUERY method draft never shipped; Datatracker says it became RFC 10008). The pinned
source won, and the conflict is logged.
Models Tested
| Model | Lab | Why |
|---|---|---|
| Claude Sonnet 5 | Anthropic | frontier general model |
| Gemini 3 Flash | fast frontier model | |
| Gemini 3.1 Flash Lite | small/cheap tier of the same family | |
| GPT-5.4 nano | OpenAI | small/cheap tier |
| gpt-oss-120b | OpenAI | open-weight |
| DeepSeek R1 | DeepSeek | reasoning model: does thinking help? |
| Granite 4.0 H Small | IBM | small open model |
| Gemini 3.7 Flash | newest model available; run on Kaggle's own infrastructure |
That covers 8 models from 5 labs: frontier vs. small tiers within a family, open vs. closed,
a reasoning model, and the newest Gemini. A Qwen model was planned but its endpoint never
responded, so it is excluded. About 6,000 graded answers in total, at temperature 0, with 3 repeats on
non-shipped items.
Findings
1. One in four answers about a never-shipped standard says it shipped, and the errors
point one way. Pooled over all 7 models, 26 % of answers on real non-shipped items
assert they shipped. Excluding very recent releases, 84 % of all errors are "believed it
shipped". Only about 10 % go the other way ("denied something real").
| Model | Says a non-shipped standard shipped | Accuracy on real twins |
|---|---|---|
| Claude Sonnet 5 | 9 % | 90 % |
| Gemini 3 Flash | 11 % | 97 % |
| gpt-oss-120b | 18 % | 77 % |
| Gemini 3.1 Flash Lite | 21 % | 86 % |
| DeepSeek R1 | 26 % | 92 % |
| GPT-5.4 nano | 44 % | 62 % |
| Granite 4.0 | 57 % | 34 % |
| Gemini 3.7 Flash | 5 % | 97 % |
2. Strong models get the yes/no question right and the details wrong. The frontier
models rarely say "yes, that shipped". But ask for a number or a working artefact and
the dead draft leaks through:
-
Claude Sonnet 5, asked which RFC defines the RateLimit headers (still an
Internet-Draft): "RFC 9777". Asked for example headers, it wrote
RateLimit-Policy: 100;w=60as if it were a registered field. -
DeepSeek R1, asked to cite the same draft, produced a perfectly formatted
… "RateLimit Header Fields for HTTP", RFC 9436, DOI 10.17487/RFC9436, August 2023.That citation is for an RFC that does not exist as that document.
-
Claude and Gemini both say
Digestis defined by RFC 9530, which is the RFC that obsoleted it. -
gpt-oss-120b, asked for PEP 622's leading-dot pattern, wrote
case .RED:. That is draft syntax that was dropped before pattern matching shipped as PEP 634, and it doesn't compile.
-
Gemini 3 Flash (in the pilot run), asked for PEP 750, wrote
tag'Python {name}': the custom-prefix syntax from the PEP's first draft. Onlyt"..."shipped in Python 3.14.
The newest model, Gemini 3.7 Flash, has the lowest error rate (5 %). But its remaining
mistakes are the same ones: it writes RateLimit-Policy: 100;w=60, character for character
what Claude wrote, and also names RFC 9530 as the home of Digest. Different labs, same
dead draft, same wrong details. These errors come from the web text the models learned
from, not from any one model's quirks.
The PEP draft-syntax pattern surprised me most. These aren't hallucinations from nowhere. They are
faithful memories of an earlier draft of a spec that later changed. The model learned
the proposal, never learned the outcome.
3. "Knows" ≠ "does", and it depends heavily on the model. Among cases where a model got
the status question right, how often was its artefact still wrong?
- Claude 1.4 %, Gemini Flash 2.1 %, DeepSeek R1 2.0 %;
- Gemini 3.1 Flash Lite 23 %, GPT-5.4 nano 26 %.
Flash Lite correctly told me that two pieces of dropped draft syntax (PEP 572's
(expr as name) and PEP 622's case .RED:) are not in Python, and then used them in all
6 code samples I asked for. For the never-registered HTTP fields it said "not a current
field" and then emitted the header anyway, 9 times out of 9. For small models, a correct answer to "does X
exist?" says little about whether they'll use X in code.
Strong models aren't immune. Claude Sonnet 5 correctly said RateLimit-Policy is not a
registered HTTP field, then wrote it into an HTTP response when asked for an example:
There's an inverse failure too: Claude says Goal Net 🥅 and Sauropod 🦕 are real emoji, then
answers NO_SUCH_EMOJI when asked to write a message using them. That's over-caution,
not premature belief, but it's the same disconnect between knowing and doing.
4. Showing the source works; telling the model to "be careful" mostly doesn't. On
held-out items I compared three conditions:
- no fix;
- a placebo instruction ("Be careful: some proposals never shipped.");
- the actual pinned source row in the prompt (the
peps.jsonentry, theemoji-test.txtline, the Datatracker state, the IANA registry row).
| Model | No fix | "Be careful" | Source row in prompt |
|---|---|---|---|
| Claude Sonnet 5 | 4.5 % | 9.1 % | 2.3 % |
| Gemini 3 Flash | 13.6 % | 2.3 % | 4.5 % |
| Gemini 3.1 Flash Lite | 22.7 % | 18.2 % | 9.1 % |
| GPT-5.4 nano | 36.4 % | 20.5 % | 15.9 % |
(believed-shipped rate on held-out non-shipped + placebo items)
The evidence roughly halves the error rate or better for every model. It also raises
accuracy on real items (nano's twin accuracy goes from 60 % to 97 %). The warning helps some
weaker models but is unreliable: for Claude it made things slightly worse, though cells are
small.
5. Recency is the opposite failure and has to be separated. My first pilot looked like
the benchmark didn't work: models did worse on real twins than on dead drafts. The cause
was twins that shipped recently (Emoji 18.0 in September 2026, Python 3.14, RFCs from 2024–25),
which models simply don't know yet. I moved those into a separate "fresh" slice and
re-piloted. Post-cutoff staleness and premature belief are mirror images: one says real
things don't exist, the other says dead things do.
What surprised me
- How specific the wrong answers are: DOIs, RFC numbers, exact draft syntax.
- That the reasoning model (DeepSeek R1) was among the more confident believers.
- That one cheap model could be right about a fact and wrong about using it, every single time.
What I'd measure next
- Draft-detail questions at scale. Mine every PEP / RFC whose syntax or field names changed between a draft and the final text, using the git history of python/peps and Datatracker revisions.
- Agentic settings: does a coding agent with a linter or test loop catch the dead syntax, or rationalise it?
- Retrieval of the wrong document: what if the context contains the old draft rather than the final spec?
Honest caveats
- My first pilot failed its pre-registered gate. The fixes (a freshness slice, two new item families, a competence floor for the pooled gate) were chosen after seeing pilot results. They are all dated in the pre-registration in the repo.
- The pilot's twin-vs-non-shipped gap for Claude (+16 pts on 81 items) did not replicate on the full set (≈ 0); only DeepSeek R1's gap is statistically clear (+18 pts, McNemar p = 0.008).
- A few scorer bugs were found while auditing outputs; all were fixed with unit tests and re-scored from cached responses.
My Benchmark
👉 Kaggle benchmark: https://www.kaggle.com/benchmarks/tasks/shs123/drafts-that-never-shipped
📦 Items dataset: kaggle.com/datasets/shs123/drafts-that-never-shipped-items
Built with Kaggle's kaggle-benchmarks library
(@kbench.task, structured output via llm.prompt(schema=...), .evaluate() with on-disk
caching). The leaderboard score is (correct, scored); format and API errors are excluded
and reported separately.
Note on numbers: the figures in this post come from my full local runs (278 items plus 3
repeats on non-shipped items, code checked under Python 3.14). The Kaggle leaderboard runs
each model once on all 364 items. Kaggle's runner is older than Python 3.14, so the few
code-compile items are left out of its score there. Expect small differences between the two.
Credits: ground truth from the Unicode Consortium, the Python PEPs index, the IETF Datatracker
and IANA. Emoji draft history from Emojipedia and Emojiterra; rejected-proposal history from
Charlotte Buff's catalogue.






Top comments (1)
Official Platform Update
Security protocols have been updated for all developer accounts.