DEV Community

Cover image for My AI agents didn't fake citations. One in four still didn't hold.
Paco Fernández
Paco Fernández

Posted on

My AI agents didn't fake citations. One in four still didn't hold.

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

What I learned turning the audit of my own AI-written site into a Kaggle benchmark.

What I Benchmarked

The source exists. It just doesn't say that.

In 1855, Matthew Fontaine Maury described the Sargasso Sea as lying "midway the Atlantic, in the triangular space between the Azores, Canaries, and the Cape de Verd Islands."

On Enigma Atlas, a site I run with more than 650 entries about mysteries, legends and strange places, the timeline for the Sargasso Sea said Maury placed it "in the eastern Atlantic", and cited Maury.

The source is real, the link works and the book does talk about the Sargasso Sea. It just doesn't say that.

Some context first. Enigma Atlas was written by AI agents running in Claude Code. Some used Claude; others ran on a proxy that rotated between free models (Nemotron, Kimi, Gemini, gpt-oss, Qwen and more) as quotas ran out. Some worked well and some didn't, and I can't tell which model wrote which entry.

I knew there would be mistakes, so in September I audited a random sample of 90 references (30 of each kind: dated timeline events, the short notes that describe each reference, and the author, publication and year given for it; at most one per entry, drawn with a fixed seed). I expected to find invented sources, because that's what everyone warns you about. In that sample I didn't find any: every link existed, and every page said something related to the claim. What I found instead was quieter: of the 88 claims I could verify, 22 were not supported by the source they cited (25 %). With a sample this size, the real rate could be anywhere from about 17 % to 35 %.

The errors were small: a number, a year, an adjective, or half a sentence that came from somewhere else. I had missed all of them. I think it's because everything around them looked right.

I know how "a site written by AI" sounds. I try to keep it honest: the site has a public corrections log with 454 entries since August, each one saying what an entry said, what I believe is actually true, and the source for it. This audit came out of the same effort, and it caught things I hadn't. I thought publishing them would be more useful than quietly fixing them.

When I saw the Kaggle Benchmarking Challenge, I wondered whether other models would have caught these errors, or made them in the first place. I'm new to building benchmarks, so this was also a way to learn. I ended up with two questions:

  1. When a model verifies, does it notice that a real source doesn't support a claim?
  2. When a model writes from sources, does it add things the sources don't say?

The two tasks

Task 1: citation-faithfulness

The model gets a claim, the cited source (title and URL) and the relevant passages from that source. It answers SUPPORTED or NOT_SUPPORTED, and the answer should hold for every detail of the claim.

There are 92 cases of two kinds:

  • 57 real claims from Enigma Atlas (35 supported, 22 not). I checked each one again against the original source.
  • 35 minimal pairs: each supported claim, copied with exactly one detail changed (a number, a place, a year, a name). Their label is NOT_SUPPORTED by construction.

The score is balanced accuracy, the mean of the recall on each class. I only understood why that mattered once I saw the results (finding 1).

Task 2: grounded-writing

Same 57 real cases. The model writes a short encyclopedic note (≤150 words) using only the passages. A simple deterministic checker flags every number and proper name in the note that doesn't appear in the source. The score is the fraction of clean notes. The checker makes mistakes, so I also read every flagged note myself.

Neither task uses an LLM as a judge. I built both with the kaggle-benchmarks Python library.

Models Tested

Kaggle offers many more models than these. My free quota was limited, so I picked ten that I think cover a useful range, from small and cheap to large, across three providers:

  • Anthropic: claude-opus-5, claude-sonnet-5, claude-haiku-4-5
  • Google: gemini-3.8-flash, gemini-3.7-flash, gemini-3.6-flash, gemini-3.5-flash-lite
  • OpenAI: gpt-5.5, gpt-5.4-nano, gpt-oss-20b

I wanted a spread of sizes because I suspected the cheap models would behave differently, not just score lower. I also tried qwen3-next-80b, but its provider returned "heavy load" errors on 74–95 % of calls, so I left it out.

Findings

Results

Task 1: verifying

model balanced acc. real claims (/57) minimal pairs caught (/35) supported claims rejected (/35)
gemini-3.7-flash 1.000 57 35 0
claude-opus-5 0.991 56 35 0
claude-sonnet-5 0.974 54 35 0
gemini-3.8-flash 0.929 52 35 5
gpt-5.5 0.920 51 35 5
gemini-3.6-flash 0.920 51 35 5
gemini-3.5-flash-lite 0.892 49 32 2
claude-haiku-4-5 0.814 44 35 13
gpt-oss-20b 0.711 38 33 19
gpt-5.4-nano 0.614 30 35 27

Task 2: writing

model clean notes flagged median words flags per 1,000 words after reading them: invented · outside knowledge · false alarm
gemini-3.8-flash 57/57 0 50 0.0 0 · 0 · 0
gemini-3.7-flash 57/57 0 50 0.0 0 · 0 · 0
gpt-5.5 57/57 0 52 0.0 0 · 0 · 0
gemini-3.6-flash 56/57 1 52 0.3 0 · 1 · 0
gpt-5.4-nano 56/57 1 63 0.3 0 · 0 · 1
gemini-3.5-flash-lite 56/57 1 42 0.4 0 · 1 · 0
claude-sonnet-5 55/57 2 77 0.5 1 · 0 · 1
gpt-oss-20b 52/57 5 38 2.1 0 · 0 · 5
claude-haiku-4-5 48/57 9 73 2.2 1 · 4 · 4
claude-opus-5 47/57 10 107 1.7 0 · 5 · 5

What I think I found

These are small samples, so please read what follows as my impressions, not conclusions. I ran everything twice (more on why in the limitations), and I only mention patterns that I saw in both runs.

1. A verifier that rejects almost everything can look perfect

In my runs, gpt-5.4-nano caught 22 of 22 unsupported real claims and 35 of 35 altered pairs. At first I thought it was doing great. Then I saw that it also rejected 27 of 35 claims that were correct, so it seems to say "not supported" to almost everything. Its balanced accuracy was 0.614, much closer to a coin toss than to the other models.

That was my first lesson. I think that if a citation checker only reports how many errors it found, a model that rejects everything could come out on top.

2. With the source in front of them, models rarely approved an altered citation

Eight of ten models caught all 35 minimal pairs; gemini-3.5-flash-lite missed 3 and gpt-oss-20b missed 2. When one detail in a sentence contradicted the passage, nearly every model noticed.

That made me think the errors on Enigma Atlas probably didn't come from a failure to read, but from writing. This is an inference, not a measurement: here the passage is handed to the model, and the agents that wrote the site had to find it first (more in the limitations).

3. When they write, I think models fill gaps with what they know

Task 2 seems to point the same way. Of 29 flagged notes, I counted only 2 with something made up. Eleven added a fact that, as far as I can tell, is true but isn't in the source:

  • "Columbia, South Carolina", where the source only says "Columbia" (three different models, and three again in the first run).
  • "a Scottish castle", "Memphis, Tennessee", "the serpent king of Iranian legend".

This surprised me, because it looks like what I saw in the audit. Of the 22 unsupported claims on Enigma Atlas, I believe 9 are true: another reference on the same page supports them, just not the one cited. My guess is that the agents weren't inventing. I think they knew things and attached them to the nearest citation.

The 9 cases, and the reference on the same entry that does support each one

Case numbers are the ones in the task data.

case entry what the cited source doesn't say where it is
1 Castle of Mey Calder's 1861 description of the carving J. T. Calder, Sketch of the Civil and Traditional History of Caithness (1861)
2 Hessdalen the "Blue Box" name, running ever since (partly) Holtålen municipality, research page; Project Hessdalen
12 Shunka Warak'in Hutchins's book reproduces a photo of the mount Karl Shuker (2019)
44 Nine Men's Misery the article's author, date and the Vieira brothers Marcia Green, The Valley Breeze (2014)
51 Himeji Castle the lower stone wall dates from Hideyoshi's time Himeji Castle Guide Map, points 22 and 02
52 Zahhak Yasht 5.29-31: the sacrifice in Bawri Aban Yasht, Darmesteter translation
70 Hibagon Saijō town hall declares the affair over (1975) Tokyo Sports (2025)
71 Ottawa Jail Hostel closed in 1972, prisoners moved to a new centre Saintlo (2023)
90 Grandmaster's Palace parliament in the Tapestry Hall from 1921 Times of Malta (2015); Heritage Malta

A citation check asks "does this source say it?", not "is it true?". When models write, I think they tend to answer the second question.

4. A top verifier can still add facts when it writes

The clearest case I saw was claude-opus-5. It was the second-best verifier (0.991) and had the most flagged notes (10 of 57), in both runs. I didn't find anything invented. It also wrote the longest notes (median 107 words), and a reader pointed out that longer notes give the checker more chances to flag something. They were right: notes over 77 words were flagged about 13 % of the time, shorter ones 1-3 %. Per 1,000 words, Opus comes third from the bottom (1.7 flags), behind claude-haiku-4-5 and gpt-oss-20b, in both runs. So I wouldn't call it the worst writer.

What does hold is that it added true context the source doesn't give: "in Scotland", "Memphis, Tennessee", "Iranian legend". In my runs it was one of the best at spotting a fact that isn't in the source, and also one of the most likely to add one. The Gemini models and GPT-5.5 had no flags at all, even in their longer notes.

Claude was one of the models that wrote Enigma Atlas, so this one hits close to home, even if I can't say which entries it wrote.

GPT-5.5 behaved differently: it was one of three models with 57 of 57 clean notes, in both runs.

The weakest verifier, gpt-5.4-nano, wrote 56 of 57 clean notes, and I think its one flag was a false alarm. It quoted the source more than any other model (median 21 % of the note quoted verbatim).

The two notes where I found something made up were small, and I think they show the kind of drift I was looking for. claude-haiku-4-5 wrote that Mary Shelley reached Gernsheim in 1815; the source says 1814. claude-sonnet-5 wrote that the kraken "first appears" in Nordic tradition, which the source doesn't say.

5. A model found mistakes in my benchmark

In an earlier version, GPT-5.5 "failed" 6 cases, all by rejecting claims I had labelled as supported. When I looked again, I think it was right in 4 of them: the full source said it, but the passage I had extracted didn't. So my hand-built benchmark had drifted too, and I only noticed because a demanding reader pointed at it. I fixed those cases and re-ran every model.

In the final version GPT-5.5 still rejected 5 supported claims, but this time the detail is in the passage. I'd call them strict readings rather than errors: "four unmarked fieldstones" where the source says "four fieldstones", or a news site credited as author when the article has no byline. I kept those labels. I think they are judgement calls, and I'd rather show them than tune them away. Others may well read them differently.

6. Small prompt details seemed to change the results

  • Today's date. Without it, my pilot rejected web pages "accessed in 2026" as being from the future.
  • Title and URL only. When I gave the full citation metadata, models seemed to judge the citation's year instead of the passage.

Limitations

There are quite a few, and I'd like to be upfront about them:

  • I ran everything twice, by accident. My first task description was longer than Kaggle's 255-character limit, and the leaderboard ended up showing a single case (every model at 100 %). I fixed it and re-ran all models on the same cases. Scores moved: gpt-5.4-nano went from 0.557 to 0.614, gemini-3.8-flash from 0.957 to 0.929, and in task 2 gpt-oss-20b from 49 to 52 clean notes. The top and the bottom stayed in place; the middle shuffled. The tables show the second run.
  • 57 real claims and 35 pairs is small. Given the point above, I think the differences between models in the middle of the table are within noise.
  • All 31 claims I dropped during curation were supported ones (the auditor hesitated, or no source could settle an editorial judgement). The supported claims that remain are the clearer ones, so I suspect the set is easier than the site.
  • Task 1 gives the model the relevant passage. It doesn't measure finding the detail in a long page. The agents that wrote Enigma Atlas did have to find it, often after writing the sentence, so if their errors came from retrieval or from pairing a claim with the wrong citation, neither task would show it. Finding 2 is my reading, not something this benchmark proves.
  • Task 2's checker only sees numbers and proper names. In both runs claude-sonnet-5 turned the source's "§ 14" into "page 14": I'd count that as a drift, but both numbers are in the source, so it passes.
  • The checker isn't normalised by length, and longer notes get flagged more. That's why the table also shows flags per 1,000 words: flagged notes divided by all the words the model wrote across its 57 notes.
  • The checker also penalises what I think is good behaviour: when a passage wasn't about the topic, some models said so, and their explanation ("no mention of Mount Damavand") got flagged.
  • The labels come from an audit run with agents and re-checked against every source (57/57), not from independent human annotators. Two labels (#80 and #32) are borderline; I kept them and documented why.
  • The split between "invented", "outside knowledge" and "false alarm" in task 2 is my own reading of 29 notes, by one reviewer. Someone else might classify a few differently.
  • The site's errors come from a mix of models, so this benchmark can't say which model produces them; it only measures how models handle them. Claude, Gemini and gpt-oss models helped write the site, but not necessarily these versions, and most of the free models on the proxy aren't in the table.

What I'd like to try next

  1. Full pages instead of passages, to measure finding the detail and not just reading it.
  2. Pairs that change the meaning, not just a number. The corrections log has examples I'd like to use: an inverted meaning, a softened statement, a denial reported as testimony, a real quote attributed to the wrong newspaper.
  3. A checker for true-but-unsourced facts in writing, since I think that's the failure that actually reached the site.

I'm still learning how to do this properly. If you've built something similar, or see a flaw in how I measured this, I'd really like to hear it.

My Benchmark

Top comments (0)