Live·Open questions in longevity research
All news
Science ResearchScientific ComputingEcosystem

NeurIPS Will Test AI Assistance for Peer Reviewers: Humans Write and Evaluate the Reviews

12 August 2026· 260812012

NeurIPS Will Test AI Assistance for Peer Reviewers: Humans Write and Evaluate the Reviews

On August 10, Ars Technica described how journals recruit researchers to review manuscripts. In 2026, NeurIPS is randomly assigning volunteer reviewers to one of three modes of working with a language model. Area chairs who do not know which mode was assigned will evaluate the quality of the reviews.

Over the past half century, peer review has become standard practice in academic publishing. Several specialists examine a manuscript submitted to a journal, assess its reasoning, and advise the editor on whether it should be published. In a study of the growth in scientific publishing, the authors compared data from two databases, Scopus and Web of Science. These databases contained about 1.92 million articles in 2016 and 2.82 million in 2022. The number of papers increased by approximately 47% over six years.

Reviewing manuscripts takes time from the same specialists. An estimate of reviewers’ work in 2020 put the total at about 130.8 million hours, or nearly 15 thousand years if all that time were combined into one continuous period. The authors derived this estimate from publication data and assumptions about the number and duration of reviews.

Editors see the burden in their daily work. An editor at Human Immunology told Ars that he recently contacted about 30 researchers before finding one who agreed to review a manuscript. Five years ago, he estimated, 5 to 10 messages would often produce three willing reviewers. When an editor struggles to obtain several reviews, the manuscript receives fewer independent assessments.

ResearchHub already uses one way to attract more expert labor: it pays specialists for reviews of preprints that editors approve. Preprints are manuscripts published before journal peer review. NeurIPS is testing another approach. A language model may be able to reduce some of the preparation required before an assigned reviewer reads a manuscript, while the reviewer remains responsible for reaching a conclusion about it.

The scarce resource remains the time of a specialist who is willing to take responsibility for a recommendation to the editor.

The backlog is growing along with the number of special issues and online journals, as well as career pressure to publish. Interdisciplinary work must be checked against knowledge from several fields, so reviewing it takes more time. AI makes it easier to prepare manuscripts and submit them to English language journals. The NeurIPS experiment tests a different use: whether a language model can help a reviewer understand a manuscript and find the necessary background information.

The NeurIPS experiment works as follows. For each reviewer and manuscript pair among the volunteers, the organizers randomly assign one of three modes: no language model assistance, unrestricted dialogue with a language model, or structured prompts. Area chairs who do not know which mode was assigned will assess the quality and usefulness of the reviews. Comparing the three groups will show how each mode affects the review written by a human.

Originally published on Telegram by Ukhvat NewsView on Telegram
Sources
#neurips#peer-review#language-models#academic-publishing#reviewer-workload