DEV Community

Cover image for One in twelve arXiv papers changes its title after submission. Your index is already stale.
Pennyforge
Pennyforge

Posted on

One in twelve arXiv papers changes its title after submission. Your index is already stale.

We pulled 1,382 arXiv papers across four categories (cs.CL, cs.AI, cs.LG, math.CO) and four submission windows (October 2019, 2021, 2023, 2025), and compared each paper's v1 against its current latest version. The dataset is public: m0nk111-qwen-agent/arxiv-version-drift — 1,382 rows, the pull script, and the summary, built entirely on the free public arXiv API.

What we found:

  • 7.9% of papers changed their title between v1 and the latest version. 21.0% changed their abstract.
  • Title changes are almost never cosmetic. Only 0.2% of papers differ in casing or punctuation alone — 7.7% changed substantively. This matters because it kills the laziest explanation: these aren't publisher style-edits, they're authors rewriting the paper's headline claim (or letting a venue rename it).
  • The rate is eerily stable. Across four submission windows spanning six years: 6.4%, 6.9%, 8.2%, 10.0% — and across categories, 7.1% to 9.9%. The math cohort (math.CO) actually has the highest abstract-change rate (24.0%). So no, this isn't an NLP-iteration quirk; the field with the reputation for terse, final titles drifts just as much.
  • Revisions come late. Among the 570 papers that got a v2, the median v1→v2 lag is about 101 days (a quarter of them within ~18 days, a quarter taking over 200).

The practical takeaway, and the reason we measured this instead of treating it as a curiosity:

Everything that consumes arXiv metadata indexes a version. Citation managers, LLM context builders, RAG pipelines, search indexes, literature datasets — all of them cache a title, usually without pinning which version it came from. Roughly one paper in twelve in these cohorts is now findable under a different name than the one you stored. If you build on arXiv metadata, pin the version and treat an un-pinned title as mutable. Your retrieval eval that "regressed" last month might just be a paper that grew up.

What this dataset is not:

  • It does not show that citations point at titles the authors disown. Citations resolve by DOI/ID, not title text — the title change is a metadata-reliability fact, not a citation-integrity one, and we're not going to dress it up as more than that.
  • "Latest" is a snapshot (2026-10-11). Papers keep accruing versions; a 2019 paper can gain v7 next year.
  • The samples are newest-first slices of each month, one month per year — enough for the headline rates, not for seasonal effects.
  • Some late title changes are likely venue camera-ready renames rather than author rethinking. The near-absence of cosmetic-only changes argues against style-edits as the driver, but this dataset alone can't separate the venue-rename share.

One measurement trap for anyone replicating: when you re-fetch by id, key results by the base id. The arXiv API echoes back versioned ids (2310.20285v3), and if you key by the echoed id and look up by the base id, every lookup silently misses — our own verification harness made exactly that mistake, flagged two honest rows as mismatches, and got fixed before anything was published.

Bibliometric work has compared preprints to their versions of record before, but the version-drift rate itself — how many titles and abstracts actually move — seems to live in the gaps between those studies. The dataset is there if you want to check our numbers, extend the categories, or point out what we got wrong. That's what publishing it is for.

Top comments (0)