Science: AI Can Produce More Research Than Humans Can Verify
Science: AI Can Produce More Research Than Humans Can Verify
On 16 July, Science Editor-in-Chief Holden Thorp described the bottleneck facing AI in science. The generation of hypotheses, analyses, and manuscripts is accelerating, while verifying their origins and reliability remains a human task.
In a Science editorial published on 16 July, Thorp proposes measuring scientific AI by two rates. The first is how many hypotheses, calculations, and papers a machine can produce in a day. The second is how many researchers can verify thoroughly enough for the conclusions to be trusted. AI agents increase the first rate, while the second is limited by how many checks people can complete.
The reason is simple. A finished paper conceals a long chain of decisions: which data the system included, which data it excluded, which metric it designated as primary, and how many alternatives it tested before obtaining a successful result. When an agent constructs this chain, reading a polished manuscript is not enough for an editor. The editor needs the raw data, code, and action log to reproduce the work.
In December 2025, the authors of a study of two open autonomous scientific systems identified four types of hidden failure: an unsuitable test, data leakage, an incorrect success metric, and the selection of a favorable result after the experiment. The final manuscript may conceal these decisions. A complete action log and the code make them visible.
In May, the authors of a preprint on fabricated bibliographic citations, meaning a study that had not yet undergone peer review, examined 111 million citations across 2,5 million papers. They estimated that at least 146 932 nonexistent citations appeared in 2025. Each such citation forces an editor, reviewer, or reader to search manually for genuine evidence supporting the claim.
A scientific agent should provide a verifiable record with every result: the original question, the data, the intermediate decisions, and the final conclusion. This record makes verification part of the research process instead of an attempt to reconstruct the work from polished prose after publication.
Dorothy Chou described the path from protein prediction to a drug: after the computational stage, laboratory validation, manufacturing, and sustained funding are still required. Thorp adds editorial verification to this sequence. In biology, an error in selecting a target, model, or biomarker can consume months of laboratory work and the funds needed for the next intervention. The speed of scientific AI is determined by the speed at which its conclusions can be verified. Journals, laboratories, and funding bodies need a reproducible chain of evidence leading to every conclusion.