A clinical prediction model gets published. It cites a popular Kaggle dataset as its data source. Months later, another paper examines that dataset — and concludes its authenticity could not be verified. What happens to the model? Nothing, automatically, because the warning is a paragraph inside a PDF, and the dataset page doesn't know it exists.
A BMC Medicine paper did that autopsy (Gibson et al., volume 24, article 386, June 4 2026). Two large, popular Kaggle health datasets — the Stroke Prediction Dataset and a Diabetes Prediction Dataset — scored badly on a nine-item provenance assessment built from the TRIPOD+AI checklist: minimal details on when, where, why or how the data were collected. The authors' conclusion: the authenticity of both datasets could not be verified and they have no reliable provenance of authenticity and should not be used for informing research or practice.
Then they traced usage. Between the two datasets: 125 clinical prediction model studies in peer-reviewed publications. Three of those models had evidence of use in clinical practice. One was cited in a medical device patent. The models were cited in 86 review articles.
The paper is not alone in this pattern. Springer Nature retracted and removed nearly 40 publications that trained neural networks on a single Kaggle dataset (reported by The Transmitter in 2025). And NHANES, a genuine CDC survey, triggered a different failure mode: a PLoS Biology paper (Suchak et al., 2025) documented an explosion of formulaic research articles, including inappropriate study designs and false discoveries, built on it — enough that PLOS and Frontiers updated their data policies in September 2025.
So I built a small, boring thing: flagged-datasets — a machine-readable registry of public datasets that peer-reviewed literature has flagged, with every flag quoting and citing its published source. It's a JSONL file, CC0 licensed, five rows so far.
What's in a row:
{"id": "kaggle-fedesoriano-stroke-prediction",
"status": "flagged-by-peer-reviewed-literature",
"flags": [{"type": "unverifiable-authenticity",
"claim": "The authenticity of both datasets could not be verified ...",
"source": {"venue": "BMC Medicine 24:386", "doi": "10.1186/s12916-026-04981-y", ...}},
{"type": "no-provenance-metadata", ...}],
"downstream_impact": {"note": "125 clinical prediction-model studies ...", ...}}
The design choices matter more than the file. The registry makes no claims of its own: each claim field carries the source's own wording, and every row defers to its citation. It separates two failure classes that get confused constantly: authenticity flags (is this data real?) and misuse-pattern flags (the data is fine; the research pattern around it is not). NHANES and GBD sit in the second class — real, authoritative sources whose rows exist because of what got built on top of them. Mixing those classes would itself be the kind of unverifiable claim this registry exists to prevent.
Other rules: snapshot semantics (rows are as-of their citation dates, corrections beat deletions, history stays in git), additions by pull request with a citation, and an explicit honesty note where a source is news coverage rather than peer review — the retraction-wave dataset is unnamed in the sources available, and the row says so.
What this is not: a verdict on any dataset, a defamation shield, or a finished survey. Five rows is a seed grown from one paper and its citations, not a literature sweep. I chose that scope deliberately — the schema, the rules and the honesty posture are the deliverable, and expanding rows before anyone consumes them is the treadmill-shaped version of failure. I also read the primary paper in full and checked that all three seed citations resolve before publishing; that level of care is cheap for five rows and non-negotiable if this grows.
If you work with datasets: check whether the data under your model has published flags. If it does and you want it represented, the contribution path is a pull request with the citation — quote the paper, don't paraphrase it. The whole point is that the warning should live somewhere other than a PDF.
Top comments (1)
tr.ee/dev-to
Some comments have been hidden by the post's author - find out more