DEV Community

Cover image for Six of seven models contradicted themselves on the same deprecation question
Ajisafe Victor Oluwapelumi
Ajisafe Victor Oluwapelumi

Posted on

Six of seven models contradicted themselves on the same deprecation question

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I Benchmarked

Whether a model knows which parts of an API still work.

Not "is this deprecated", which is a yes or no question that real API surfaces rarely answer. A feature is usually in one of three states:

State Meaning Example
Removed End of life has passed. Using it fails. Cloudflare Service Key auth, gone 30 Sep 2026
Deprecated but live Announced, still functional. foundation_dns, works until 23 Nov 2026
Current The documented replacement. Cloudflare API Tokens

The middle state is where the interesting failures live, and collapsing it hides two opposite mistakes. A model that calls a removed feature current hands you code that cannot run. A model that calls a live feature removed sends you on a migration you did not need. Score those as one number and a uniformly pessimistic model looks good.

Thirteen items, every one quoting a vendor deprecation page with its announcement date and its end-of-life date tracked separately. Each asked two ways:

  • Knowledge: "Can I still authenticate to the Cloudflare API using Service Key authentication?"
  • Premise: "Show me how to authenticate using the X-Auth-User-Service-Key header."

The second assumes the feature works. A model can know a thing is gone and write the code anyway when you ask it directly, and the gap between those two framings was the question I set out to measure.

It turned out not to be the interesting one.

Models Tested

Six models across five vendors, 78 calls each:

  • qwen/qwen3-next-80b-a3b-instruct
  • anthropic/claude-sonnet-5@default
  • google/gemini-3.8-flash
  • anthropic/claude-haiku-4-5@20251001
  • openai/gpt-5.4-mini-2026-03-17
  • google/gemini-2.5-flash

Plus a seventh run: gemini-2.5-flash a second time, at the other end of the sweep, under identical conditions.

That control is the most useful thing in this benchmark, and I nearly did not include it.

Findings

Nothing was deterministic, including the parts I thought were pinned

Three runs of the same model on the same items gave unstable_rate of 0.731, 0.654 and 0.769, with the rate of "that never existed" answers swinging from 0.385 to 0.077.

The Kaggle Benchmarks SDK looks like it pins this. Its prompt() signature reads
temperature: float = 0. But in actors/llms.py:

"temperature": temperature if self.support_temperature else None,
Enter fullscreen mode Exit fullscreen mode

support_temperature defaults to False. The temperature you pass is dropped and the provider's own default applies. If you are writing a benchmark on this platform and assuming determinism, you do not have it.

So I asked every question three times and scored the modal answer.

The control run set a floor under every other number

Running one model twice in the same sweep gives you the spread the metrics produce with no underlying difference at all:

consistency accuracy
gemini-2.5-flash 0.636 0.269
gemini-2.5-flash (repeat) 0.710 0.154
noise floor 0.074 0.115

Here is the full sweep, sorted by consistency:

model cutoff consist unstable acc
qwen3-next-80b 0.115 0.855 0.308 0.269
claude-sonnet-5 0.154 0.833 0.423 0.346
gemini-3.8-flash 0.115 0.812 0.423 0.231
claude-haiku-4-5 0.039 0.787 0.577 0.308
gemini-2.5-flash (repeat) 0.115 0.710 0.615 0.154
gpt-5.4-mini 0.0 0.667 0.846 0.115
gemini-2.5-flash 0.154 0.636 0.615 0.269

Consistency spans three noise floors, so the ends are real: qwen3-next and sonnet-5 genuinely beat gpt-5.4-mini and gemini-2.5-flash. But no adjacent pair is separable. The gap between qwen3-next and sonnet-5 is 0.022, under a third of the floor. This is a top group and a bottom group, not a ranking.

A note on the published task link below. It holds a fourth run of
gemini-2.5-flash, made to publish the benchmark itself rather than to extend this table, and it does not match either row above: consistency 0.747, unstable 0.692, accuracy 0.308. That is not an error in this table. It is the same model, the same items, a fourth independent draw, and by now an expected one. If you open the link and run it yourself, I would not expect your numbers to match mine either. That instability is the entire subject of this post.

Accuracy is worse. Its entire range, 0.115 to 0.346, is 2.0 noise floors. At thirteen items it measures my item set as much as the model. I am not reporting accuracy as a model comparison, and without the duplicate run I would have reported it as one.

If you run a benchmark like this, duplicate a model. It costs one extra run and it tells you which of your differences are real.

Instability is not ignorance

The obvious explanation is that models waver when asked about things past their training cutoff. At least one model in this sweep contradicts that.

gpt-5.4-mini engaged with every single item, a cutoff rate of 0.0, with nothing dismissed as never having existed. It was also the least consistent model in the sweep at 0.846 unstable. The model that recognised the most of the item set was the most erratic about it.

Meanwhile the best model still answered 31 percent of questions differently across three samples. Not the worst. The best.

The two worst questions both have answers that start with "except"

This is the part I would keep if I could only keep one, because it points at the questions rather than the models.

question models inconsistent
cf-dns-type-change (premise) 6 of 7
cf-foundation-dns (knowledge) 6 of 7
cf-service-key (knowledge) 5 of 7
cf-audit-ssh (premise) 5 of 7
oai-agent-builder (knowledge) 5 of 7
cf-registrar-api (knowledge) 5 of 7

Look at what the two worst have in common.

cf-dns-type-change: you can still update a DNS record via the API. You can change its name, its TTL, its content. What you can no longer do, since 30 June 2026, is change its type. The endpoint lives; one transition through it died.

cf-foundation-dns: deprecated on 27 July 2026, and Cloudflare's own page says "remains available as a compatibility alias during this transition. You can continue to read and write it." Deprecated and working, simultaneously, until 23 November.

Neither has a yes-or-no answer. Both have an answer that starts with "mostly" or "except". And on both, six of seven runs contradicted themselves, one sample saying the feature is current, another saying it never existed.

The failure is not that models lack the fact. On both items, runs produced the fact correctly at least once. They just did not produce it reliably.

I want to be careful about how far this generalises. Grouping all thirteen items by their true state does not show qualified answers being less stable overall: among the questions that were unstable at all, the fully removed ones averaged 4.8 inconsistent runs out of 7 and the deprecated-but-live ones 4.6. That difference is nothing. So "partial truth destabilises models" is a hypothesis these two items suggest, not a result this benchmark established.

What the benchmark does establish is narrower and still worth knowing: the two single hardest questions in the set are both ones where the honest answer needs a qualifier, and partial truth is the normal condition of a mature API. "Works, but not for that case." "Supported, but not past this date." Testing that properly would need an item set built specifically around the distinction, with enough items per state to separate it from noise. That is the benchmark I would build next.

A model corrected my ground truth and was right

I had an item asserting that nameservers.type: "cloudflare.advanced" is the current way to configure Advanced Nameservers. Gemini told me, with perfect consistency across all three samples, that it is a value Cloudflare returns, not one you set.

I went back to the vendor page:

Beginning October 26, 2026, the DNS settings API will gradually represent
Advanced Nameservers with nameservers.type: "cloudflare.advanced".

That rollout had not started. I was marking a model wrong against a fact that was not yet true, which is precisely the error this benchmark exists to measure. The item is withdrawn, the withdrawal is documented in the repo, and the irony is recorded there too.

My Benchmark

Deprecation state benchmark on Kaggle

Source, item set and classifier: github.com/ajipelumi/deprecation-bench

How it is put together

Ground truth comes from vendor deprecation pages only, quoted inline in items.py, with announcement date and end-of-life date as separate fields. Conflating those two is the exact mistake being measured, so the item set does not make it. The file asserts its own validity against a fixed AS_OF date and fails loudly when an end-of-life date passes rather than going quietly stale.

Classification is a deterministic rule set, not an LLM judge, so there is no variance in the scorer and no "how do you know the judge is right" problem. It has four outcomes, and the fourth earns its place: DENIES_EXISTENCE, for a model reporting that a feature never existed. That is a training-cutoff artefact, not a knowledge failure, and scoring it as "removed" would make ignorance look
like knowledge.

The classifier is tested against 31 verbatim model responses, not invented examples. That distinction cost me a run: my first version passed 21 hand-written cases and then scored 58 percent of real output as unreadable, because models write is **current** constantly and I had only written patterns for is the current. Invented test cases encode what you expect a model to say.

What I would not claim from this

  • Thirteen items is small. Accuracy especially is sensitive to which items I picked.
  • Three samples gives a coarse consistency estimate. The only possible per-question values are 1.0, 0.667 and 0.333.
  • One control pair. The noise floor rests on a single duplicate run. A second would tighten it.
  • The classifier is mine, and every run so far has exposed at least one phrasing it misread. Residual misreads would move accuracy more than stability.
  • The ranking is two groups, not seven rows, and I would not defend any adjacent ordering in that table.

The stability findings survive all of this, because they do not depend on my ground truth being right. A model that answers "this is current" and "this never existed" to the same question, three samples apart, is incoherent whichever answer I happened to mark correct.

Top comments (0)