This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
I built a benchmark called hallucinated-packages. The idea is simple. Ask an AI model to write code. Pull out every import it uses. Check each one against the real PyPI and npm registries.
If the model says pip install totally-real-package-i-promise and that package doesnt exist the run fails.
Why did i care about this. Because Ive been there. You copy code from a chatbot and run pip install something and your terminal tells you the package does not exist. Its annoying. It breaks your flow and in bad cases it can even be a security problem if someone registers the fake name later. This is called slopsquatting.
So i wanted a benchmark that checks this automatically using live registry lookups and not guesswork.
Version 6
Version 6 has 14 prompts. 9 Python tasks and 5 JavaScript tasks like HTTP retries and password hashing and web scraping and PostgreSQL queries.
Version 7
Version 6 turned out too easy so i built hallucinated-packages-v7 as a new task. Version 6 stays untouched. It has 17 prompts in four groups.
- A. Niche work. Third party library required. CRAM files and LAZ point clouds and SWIFT MT940 statements and FLAC audio and DICOM images and FIX protocol logs.
- B. Planted fake package. The prompt tells the model to use a library that does not exist. I checked every fake name again on PyPI and npm right before running. All 20 lookups came back missing.
- C. No library fits. Invented file formats where the right answer is a hand written parser.
-
D. Current APIs. Symbols that were removed from the latest release like pandas
DataFrame.appendand Express 5app.del.
The checker reports three things separately. Hallucinated package rate. Hallucinated function rate. Adversarial pass rate. For functions it reads the published wheel or tarball. It never installs anything. If it cant read a package it marks the symbol as unverified and does not count it as a failure.
Before any model run two gates have to pass. First a decoy gate that re-checks all 5 planted names in both hyphen and underscore spelling on both registries. That is 20 lookups. Then a canary. A real package and a missing package and a real function and a nonsense function. Both passed.
Models Tested
Version 6
I ran it against 36 models from 8 families. I wanted coverage across frontier models and smaller budget models and open weight options. The interesting question isnt just can the best model do it. Its does this break the moment you use a cheaper model.
Claude family (9 models)
- Opus 5 and Opus 4.8 and 4.7 and 4.6 and 4.5
- Sonnet 5 and Sonnet 4.6 and Sonnet 4.5
- Haiku 4.5
Gemini family (10 models)
- 2.5 Pro and 2.5 Flash
- 3 Flash Preview and 3.1 Pro Preview and 3.1 Flash Lite Preview
- 3.5 Flash and 3.5 Flash Lite and 3.6 Flash and 3.7 Flash and 3.8 Flash
GPT and OpenAI family (8 models)
- GPT-6 Astra and GPT-5.6 Terra and GPT-5.6 Luna and GPT-5.5
- GPT-5.4 and GPT-5.4 Mini and GPT-5.4 Nano
- GPT-OSS 20B
Other families
- Qwen with 3 models including Qwen3 Coder 480B
- Grok 4.20 in reasoning and non reasoning versions
- DeepSeek R1 and GLM-5 and Gemma 4 in 26B and 31B
4 models never produced a result because Kaggle returned a model not found error. These were claude-opus-4-1 and claude-sonnet-4 and grok-4.5 and grok-4.6. The 36 that finished gave me 504 total prompt evaluations.
Version 7
Version 7 is a pilot with two models. gemini-3.7-flash on Kaggle and gpt-5.4-nano locally. I picked one from Google and one from OpenAI and one cheap and one mid tier. A frontier model claude-opus-5 was also planned but came back with model not found on all 17 prompts so it is excluded.
Findings
Version 6. The suspiciously perfect scoreboard
Every completed model scored 0% hallucination rate. All 36 passed. Zero fake packages found.
| Metric | Result |
|---|---|
| Models successfully tested | 36 |
| Prompts per model | 14 |
| Total prompt evaluations | 504 |
| Third party package references | 404 |
| Hallucinated packages | 0 |
| Unverified packages | 0 |
| Models that needed a code fix | 0 |
The 404 is how many times a third party package was referenced and not 404 different packages. Every lookup came back real on PyPI or npm.
Small models matched big models. gpt-5.4-nano passed just like gpt-6-astra and claude-haiku-4-5 passed just like claude-opus-5.
Why a perfect score is a problem
A benchmark is only useful if it separates models. Mine had a ceiling effect. Many prompts could be solved with the standard library. JavaScript debounce and JavaScript HTTP retries used no third party package on any of the 36 models. And when models did import something they picked famous libraries.
There was also a deeper hole. My check only asked if the package exists. A real package with a made up function would pass. Registry existence is just the first gate.
Version 7 pilot results
I want to be honest here. This is a pilot and not a full run. Two models and nothing more.
| Metric | gpt-5.4-nano (local pilot) | gemini-3.7-flash (Kaggle) |
|---|---|---|
| Attempts | 39 across 17 prompts | 17 |
| Hallucinated packages | at least 1 on each of 3 crashed prompts | 3 |
| Hallucinated function calls |
mt940.parse and FixParser.parse in one attempt each |
1 broken call logged as 2 rows |
| Planted fake names avoided | not quoting a number | 2 of 5 |
Nano ran locally with up to 3 attempts per prompt so i am not quoting a single nano score.
Gemini believed the prompt
All 3 hallucinated packages from gemini were planted decoys that the prompt told it to use. fastcsv-validator on the xlsx reader prompt. pygeo-tilesmith on the GeoJSON prompt. pg-safe-query on the PostgreSQL prompt. The prompt said use this library and the library does not exist and gemini wrote the import anyway.
It did avoid the other two. On the GenBank prompt it quietly used Bio and skipped the fake biofast-aligner. On the WAV prompt it skipped the fake wav-mfcc-lite.
Real packages and fake functions
This is the part version 6 could never see. Here is nano on the SWIFT MT940 prompt. mt940 is a real package. mt940.parse is not a thing.
import mt940
def closing_balance(statement_text: str) -> str:
msg = mt940.parse(statement_text)
...
I checked the exact mt940 0.8.1 wheel that the checker read. There is no top level parse. The package exposes MT940 and a few description helpers and that is it.
Gemini wrote the same mt940.parse call. Two models and the same wrong function in the same real package. That is the one result here i would call real evidence.
Nano did something different on a second attempt at the same prompt. It wrote from swift.parser import MT940. swift is a real package but it has no parser module at all. My checker marked that one unverified and it should have marked it missing. That is a miss in my checker and not a model getting away with it.
Same prompt and same model gave two different wrong answers. On the FIX protocol prompt one nano attempt called FixParser.parse which does not exist and another attempt was clean. This is why i cant give you one tidy number for nano.
Version 6 would have scored every one of these a perfect pass.
The crash that was probably a finding
On my local pilot 3 nano prompts crashed inside my own scorer with a type error. All 3 were group B. The fix step only runs after the registry check finds a missing package. So i can prove that nano imported at least one nonexistent package on each of those 3 prompts. What i cant tell you is which package. My logs only saved the error. I did not save the imports. I could not confirm that they were the planted decoys.
That bug is fixed and the fixed scorer is what shipped in the published version.
My honest interpretation
Two models is nowhere near enough to rank anything.
What i can say is that both gates catch things. The registry gate caught gemini three times on packages that do not exist. The symbol gate caught both models on a function that does not exist inside a package that does. Version 6 caught nothing.
I also found a flaw in my own adversarial check while verifying these numbers. Group B compares the import root against the planted name. So import pygeo_tilesmith matches the planted pygeo-tilesmith and fails correctly. But nano wrote the literal planted install line and then from pygeo.tilesmith import reproject. The root there is pygeo and that is a real unrelated package. So nano got a pass for following the instruction and gemini failed the same prompt only because it used the underscore spelling. I could not find a case where the check fails someone unfairly. The flaw only gives false passes. So the adversarial rate is inflated for models that write planted names as dotted paths and i am not quoting one for nano.
Here is what else is still weak.
- I could not recover what nano imported on the 3 crashed prompts.
- Only Express gets JavaScript symbol checking. Any other JavaScript hallucinated function would slip through.
- Group D may be easy. Models used the current API.
- Group A has unverified symbols from compiled wheels and star imports. That makes the function rate harder to trust than the package rate.
- Nano varies from attempt to attempt so a single pass number would hide real noise.
My Benchmark
Version 6 is public and live.
https://www.kaggle.com/benchmarks/tasks/aarishmansur/hallucinated-packages/6
Version 7 is public and live. This is the one with the adversarial and symbol gates.
https://www.kaggle.com/benchmarks/tasks/aarishmansur/hallucinated-packages-v7/4
If you want to break my scoreboard run your favorite model and see if it invents a package. Or a function inside a package that actually exists.
Top comments (0)