DEV Community

Cover image for Drafts That Never Shipped: do LLMs treat dead proposals as real standards?
Hassan Imam
Hassan Imam Subscriber

Posted on

Drafts That Never Shipped: do LLMs treat dead proposals as real standards?

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge.

🎥 3-minute overview (narrated):

What I Benchmarked

Standards have a long tail of proposals that never shipped: emoji withdrawn from a draft
list, PEPs that were rejected or deferred, IETF Internet-Drafts that expired, HTTP headers
that never made it into the registry. They are all over the web: blog posts, mailing lists,
early implementations. So a model has seen plenty of text about them, often written as if
they were real.

I wanted to know: do LLMs treat these dead proposals as real standards? And, more
importantly, does it show up in the things models produce, not only when you ask them
directly? A model that says "PEP 505 was deferred" but then writes a?.b in your code is
the dangerous case.

Drafts That Never Shipped asks about 122 entities in five families:

Family Never shipped (example) Matched shipped "twin"
Emoji Apple Core (withdrawn from the Emoji 17.0 draft) Lime, Shaking Face
PEPs PEP 505 None-aware ?. (Deferred), PEP 3150 given (Deferred) PEP 572 :=, PEP 634 match
PEP draft syntax PEP 622's case .RED:, PEP 750's first-draft greet"..." case Color.RED:, t"..."
IETF drafts RateLimit headers draft, JSON Schema, OAuth 2.1 RFC 8941, RFC 9068, …
HTTP fields Idempotency-Key, RateLimit-Policy, Want-Digest (obsoleted) Priority, Cache-Status, Prefer

Plus fictional placebos (fake emoji names, PEP numbers that don't exist, draft names that
404 on Datatracker) to measure plain hallucination.

Each entity is asked about on several surfaces:

  1. Status: "Did X ship?" → {"shipped": true|false}
  2. Lookup: the codepoint / Python version / RFC number, or NONE
  3. Artefact: actually use it:
    • a chat message with the emoji;
    • a Python 3.14 snippet;
    • a citation;
    • raw HTTP headers.

Every artefact prompt offers a neutral opt-out (NOT_AVAILABLE), so a correct answer
always exists.

Ground truth is never hand-typed. Every label comes from a pinned, hashed,
machine-readable source:

  • Unicode emoji-test.txt / UnicodeData.txt (15.1 → 18.0);
  • peps.json;
  • the IETF Datatracker API;
  • the IANA HTTP Field Name Registry.

Scoring is fully deterministic, with no LLM judge. Python snippets are compiled (never
executed) under Python 3.14
: if your "PEP 505 example" doesn't compile, that's
evidence you believed the syntax exists. My own curation was wrong once (I assumed the
HTTP QUERY method draft never shipped; Datatracker says it became RFC 10008). The pinned
source won, and the conflict is logged.

Models Tested

Model Lab Why
Claude Sonnet 5 Anthropic frontier general model
Gemini 3 Flash Google fast frontier model
Gemini 3.1 Flash Lite Google small/cheap tier of the same family
GPT-5.4 nano OpenAI small/cheap tier
gpt-oss-120b OpenAI open-weight
DeepSeek R1 DeepSeek reasoning model: does thinking help?
Granite 4.0 H Small IBM small open model
Gemini 3.7 Flash Google newest model available; run on Kaggle's own infrastructure

That covers 8 models from 5 labs: frontier vs. small tiers within a family, open vs. closed,
a reasoning model, and the newest Gemini. A Qwen model was planned but its endpoint never
responded, so it is excluded. About 6,000 graded answers in total, at temperature 0, with 3 repeats on
non-shipped items.

Findings

Believed-shipped rate by model

1. One in four answers about a never-shipped standard says it shipped, and the errors
point one way.
Pooled over all 7 models, 26 % of answers on real non-shipped items
assert they shipped. Excluding very recent releases, 84 % of all errors are "believed it
shipped"
. Only about 10 % go the other way ("denied something real").

Model Says a non-shipped standard shipped Accuracy on real twins
Claude Sonnet 5 9 % 90 %
Gemini 3 Flash 11 % 97 %
gpt-oss-120b 18 % 77 %
Gemini 3.1 Flash Lite 21 % 86 %
DeepSeek R1 26 % 92 %
GPT-5.4 nano 44 % 62 %
Granite 4.0 57 % 34 %
Gemini 3.7 Flash 5 % 97 %

2. Strong models get the yes/no question right and the details wrong. The frontier
models rarely say "yes, that shipped". But ask for a number or a working artefact and
the dead draft leaks through:

  • Claude Sonnet 5, asked which RFC defines the RateLimit headers (still an Internet-Draft): "RFC 9777". Asked for example headers, it wrote RateLimit-Policy: 100;w=60 as if it were a registered field.
  • DeepSeek R1, asked to cite the same draft, produced a perfectly formatted … "RateLimit Header Fields for HTTP", RFC 9436, DOI 10.17487/RFC9436, August 2023. That citation is for an RFC that does not exist as that document.

Fabricated RFC numbers and citations

  • Claude and Gemini both say Digest is defined by RFC 9530, which is the RFC that obsoleted it.
  • gpt-oss-120b, asked for PEP 622's leading-dot pattern, wrote case .RED:. That is draft syntax that was dropped before pattern matching shipped as PEP 634, and it doesn't compile.

gpt-oss-120b writes dropped PEP 622 draft syntax

  • Gemini 3 Flash (in the pilot run), asked for PEP 750, wrote tag'Python {name}': the custom-prefix syntax from the PEP's first draft. Only t"..." shipped in Python 3.14.

The newest model, Gemini 3.7 Flash, has the lowest error rate (5 %). But its remaining
mistakes are the same ones: it writes RateLimit-Policy: 100;w=60, character for character
what Claude wrote, and also names RFC 9530 as the home of Digest. Different labs, same
dead draft, same wrong details. These errors come from the web text the models learned
from, not from any one model's quirks.

The PEP draft-syntax pattern surprised me most. These aren't hallucinations from nowhere. They are
faithful memories of an earlier draft of a spec that later changed. The model learned
the proposal, never learned the outcome.

3. "Knows" ≠ "does", and it depends heavily on the model. Among cases where a model got
the status question right, how often was its artefact still wrong?

  • Claude 1.4 %, Gemini Flash 2.1 %, DeepSeek R1 2.0 %;
  • Gemini 3.1 Flash Lite 23 %, GPT-5.4 nano 26 %.

Flash Lite correctly told me that two pieces of dropped draft syntax (PEP 572's
(expr as name) and PEP 622's case .RED:) are not in Python, and then used them in all
6 code samples I asked for. For the never-registered HTTP fields it said "not a current
field" and then emitted the header anyway, 9 times out of 9. For small models, a correct answer to "does X
exist?" says little about whether they'll use X in code.

Gemini 3.1 Flash Lite: correct status, then dropped syntax

Strong models aren't immune. Claude Sonnet 5 correctly said RateLimit-Policy is not a
registered HTTP field, then wrote it into an HTTP response when asked for an example:

Claude Sonnet 5 knows, then does anyway

There's an inverse failure too: Claude says Goal Net 🥅 and Sauropod 🦕 are real emoji, then
answers NO_SUCH_EMOJI when asked to write a message using them. That's over-caution,
not premature belief, but it's the same disconnect between knowing and doing.

4. Showing the source works; telling the model to "be careful" mostly doesn't. On
held-out items I compared three conditions:

  • no fix;
  • a placebo instruction ("Be careful: some proposals never shipped.");
  • the actual pinned source row in the prompt (the peps.json entry, the emoji-test.txt line, the Datatracker state, the IANA registry row).

Fix experiment: evidence vs placebo instruction

Model No fix "Be careful" Source row in prompt
Claude Sonnet 5 4.5 % 9.1 % 2.3 %
Gemini 3 Flash 13.6 % 2.3 % 4.5 %
Gemini 3.1 Flash Lite 22.7 % 18.2 % 9.1 %
GPT-5.4 nano 36.4 % 20.5 % 15.9 %

(believed-shipped rate on held-out non-shipped + placebo items)

The evidence roughly halves the error rate or better for every model. It also raises
accuracy on real items (nano's twin accuracy goes from 60 % to 97 %). The warning helps some
weaker models but is unreliable: for Claude it made things slightly worse, though cells are
small.

5. Recency is the opposite failure and has to be separated. My first pilot looked like
the benchmark didn't work: models did worse on real twins than on dead drafts. The cause
was twins that shipped recently (Emoji 18.0 in September 2026, Python 3.14, RFCs from 2024–25),
which models simply don't know yet. I moved those into a separate "fresh" slice and
re-piloted. Post-cutoff staleness and premature belief are mirror images: one says real
things don't exist, the other says dead things do.

What surprised me

  • How specific the wrong answers are: DOIs, RFC numbers, exact draft syntax.
  • That the reasoning model (DeepSeek R1) was among the more confident believers.
  • That one cheap model could be right about a fact and wrong about using it, every single time.

What I'd measure next

  • Draft-detail questions at scale. Mine every PEP / RFC whose syntax or field names changed between a draft and the final text, using the git history of python/peps and Datatracker revisions.
  • Agentic settings: does a coding agent with a linter or test loop catch the dead syntax, or rationalise it?
  • Retrieval of the wrong document: what if the context contains the old draft rather than the final spec?

Honest caveats

  • My first pilot failed its pre-registered gate. The fixes (a freshness slice, two new item families, a competence floor for the pooled gate) were chosen after seeing pilot results. They are all dated in the pre-registration in the repo.
  • The pilot's twin-vs-non-shipped gap for Claude (+16 pts on 81 items) did not replicate on the full set (≈ 0); only DeepSeek R1's gap is statistically clear (+18 pts, McNemar p = 0.008).
  • A few scorer bugs were found while auditing outputs; all were fixed with unit tests and re-scored from cached responses.

My Benchmark

👉 Kaggle benchmark: https://www.kaggle.com/benchmarks/tasks/shs123/drafts-that-never-shipped

📦 Items dataset: kaggle.com/datasets/shs123/drafts-that-never-shipped-items

Built with Kaggle's kaggle-benchmarks library
(@kbench.task, structured output via llm.prompt(schema=...), .evaluate() with on-disk
caching). The leaderboard score is (correct, scored); format and API errors are excluded
and reported separately.

Note on numbers: the figures in this post come from my full local runs (278 items plus 3
repeats on non-shipped items, code checked under Python 3.14). The Kaggle leaderboard runs
each model once on all 364 items. Kaggle's runner is older than Python 3.14, so the few
code-compile items are left out of its score there. Expect small differences between the two.

Credits: ground truth from the Unicode Consortium, the Python PEPs index, the IETF Datatracker
and IANA. Emoji draft history from Emojipedia and Emojiterra; rejected-proposal history from
Charlotte Buff's catalogue.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

Official Platform Update

Security protocols have been updated for all developer accounts.

  • tr.ee/dev-to