A chain of language models went from 11 083 papers to a computational physics manuscript and tested itself by reproducing published calculations
A chain of language models went from 11 083 papers to a computational physics manuscript and tested itself by reproducing published calculations
Haonan Huang, a physicist at Princeton University, described an autonomous research system that selects a topic from a recent arXiv corpus, validates its methods against published studies, runs new calculations, and assembles a manuscript. The work was released as an arXiv preprint and accepted at the ICML 2026 AI for Science workshop. Its status helps put the result in context: this is an early scientific report, and the physical conclusion still requires independent verification.
It is easy to reduce this story to a superficial claim: AI wrote a scientific paper. That claim tells us little. A language model can produce plausible text, cite papers, and sound confident even when its conclusion has no real basis.
In the preprint from July 2, Huang tries to build a different kind of system. In this run, the agent began with 11 083 recent papers on condensed matter physics, selected a research direction in altermagnetic piezomagnetism, and proceeded through six stages to produce a manuscript. Here, altermagnetism can be understood as a particular type of magnetic order: the material retains zero net magnetization overall, while its electronic structure still contains an ordered spin asymmetry.
Before performing new calculations, the agent had to reproduce published results: it took studies by other physicists, reran their calculations, and compared its numbers with those reported in the literature. This turns a citation from decoration into a test. The statement "according to paper X" now carries a cost: the agent's calculation must withstand comparison with an external numerical result.
The report calls this mechanism grounding through anchors. In total, the system completed 47 separate clean context sessions. Each session began without any conversation history and communicated with the other stages only through files on disk. The process included 2 162 literature interactions and separate stages for topic selection, tool validation, reproduction, new calculations, and peer review. The accompanying Zenodo archive contains the prompts, rules, list of 11 083 arXiv papers, source interaction logs, verdicts for five reproduction attempts, and a record of human interventions.
That intervention record is essential to the claim. The author states that the human provided operational support: obtaining access to paywalled papers, authorizing the system to continue after failures, and adding general rules to the knowledge base after unsuccessful reproduction attempts. A table attributes the choice of research direction, parameters, and interpretation to the agent chain. The claim therefore depends on the limited role of the human as a technician and curator of rules, while scientific decisions remain with the agent.
The study also reveals its own weak point. One published numerical "anchor" used for comparison later appeared to be unstable when the author investigated it further: the value depended on the choice of computational window and basis. This provides a useful warning about the entire approach. Agreement with the literature establishes agreement with the literature. Establishing truth also requires validation of the anchor itself.
A similar validation failure has already appeared in virtual cell tests: if the metric rewards an average response, the model learns to perform well on the test while losing biological meaning. In physics, Huang tries to give the agent a more demanding test: match a published number, identify the source of any discrepancy, and preserve the calculation record.
A poor scientific agent sounds convincing and conceals errors inside polished prose. A stronger architecture exposes its calculations, failed reproduction attempts, warnings from review sessions, and every point where its result diverges from a published number. For aging research, biology, and drug discovery, this distinction will be decisive: an autonomous researcher becomes useful when it can be made to confront reality through verifiable measurements.