DEV Community

Breach Protocol
Breach Protocol

Posted on • Originally published at groundtruth.day

NeurIPS is running a randomized experiment on AI-assisted review

NeurIPS 2026 is running a voluntary, randomized experiment in AI-assisted peer review: participating reviewers are assigned, per eligible paper, to no LLM assistance, open-ended LLM assistance, or structured LLM assistance, through an interface built into OpenReview. The conference simultaneously bans unsanctioned model use in reviewing and prohibits prompt injection aimed at swaying reviews. It is the largest controlled test yet of whether machine assistance improves or degrades scientific peer review.

Key facts

  • Reviewers are randomly assigned per paper to one of three conditions - no assistance, open-ended assistance, or structured assistance - inside OpenReview.
  • The Evaluations and Datasets Track bans LLMs entirely in review; serious violations can escalate to desk rejection of the reviewer's own submissions.
  • A separate PNAS study of 7.3 million articles estimates 57% of 2025 papers showed lexical evidence of LLM influence, up from 12% in 2023.
  • Primary source: the NeurIPS AI-Assisted Reviewing Experiment page and the Main Track Handbook.

Peer review is the mechanism by which scientific claims earn the right to be taken seriously. Someone qualified reads your work, tries to break it, and reports whether it holds. At a conference receiving tens of thousands of submissions, that mechanism runs on unpaid volunteers under deadline - which is precisely the condition under which people reach for tools that make the job faster.

NeurIPS's response is unusually disciplined: rather than banning the tools and hoping, or permitting them and shrugging, it is running an experiment designed to produce evidence. Randomize the assistance condition, hold everything else constant, and measure what happens to review quality. That is how you would answer the question if you actually wanted to know.

The rules differ by track, and most of the confusion in the current discourse comes from collapsing them. Main Track reviewers may use an LLM only inside the experiment, and must use the sanctioned one rather than a model of their choice. Reviewers in the Evaluations and Datasets Track may not use any LLM or agent, and low-quality reviews may be investigated for inappropriate use, with serious violations escalating to area chairs and possible desk rejection of the reviewer's own submissions. The Position Paper Track is stricter still: papers must be substantially human-written with AI limited to copy-editing, and organizers have already used detection tooling to desk-reject some submissions.

So "AI-generated reviews at NeurIPS" is ambiguous unless someone specifies whether the review came from the sanctioned experiment or from a rule violation. That ambiguity is doing enormous work in the current round of accusations.

And there are accusations. During rebuttal week, reviewers reported submissions whose papers and rebuttals both appeared wholly machine-written. Others reported reviews and meta-reviews that read as generated. Multiple participants reported finding the same hidden instruction inside PDFs downloaded from OpenReview but absent from their own submitted files - text demanding that any LLM's review include three specific stock phrases - and inferred that NeurIPS had planted a tripwire to catch reviewers uploading submissions to outside models. A further claim held that this tripwire was routing papers to the ethics committee.

None of that is organizer-confirmed. The handbook prohibits author prompt injections and acknowledges a gray zone of prose written to appeal to machine readers, but announces no conference-side tripwire. Official guidance tells reviewers who find hidden instructions to report them to area chairs, and ethics referral is a separate process for flagged ethics concerns. One circulating claim is flatly contradicted by the handbook: authors can see and respond to ethics-review comments.

Sitting behind all of it is a number being badly misused. Kyle Siler's PNAS paper, The diffusion of large language models in published academic articles, analyzed the full text of 7.3 million journal articles published from 2020 to 2025 across four publishers - Elsevier, Frontiers, MDPI and PLOS. It estimates that 57% of 2025 articles showed evidence of LLM influence, up from 12% in 2023.

Two corrections are essential. First, this is not "over half of all academic articles" - it is four publishers, not a census. Second, and more important, "LLM influence" is a lexical proxy: the study built a set of 228 focal words whose frequency jumped after 2022 in ways consistent with model output, then scored articles on their usage. It is a smoke detector, not a fingerprint. It cannot identify which model was used, how much of a paper was generated, whether the use was disclosed, or whether anything improper happened. The paper itself says the range runs from subtle linguistic influence to mostly generated text. See perplexity for why detecting machine text from word statistics alone is so slippery.

The strongest counter-argument to the whole panic is practical, and it came from within the reviewer threads: a competent reviewer using a model may still produce a better review than a disengaged human, and penalizing polished prose falls hardest on researchers who do not write English natively. That is exactly the question the randomized experiment was built to answer, which is more intellectually honest than most of the commentary surrounding it.

The honest caveat is that a randomized trial cannot fix a trust problem. Reviewers suspect authors, authors suspect reviewers, and both suspect the conference - and a rumor about a secret tripwire spreading unchallenged is itself a measure of how little of that trust remains. Evidence will help. It will arrive after this cycle's decisions.

See also: NeurIPS bans prompt injection in reviews and LLM as a judge.


Originally published on Ground Truth, where every claim is checked against the primary source.

Top comments (0)