Live·Open questions in longevity research
All news
Scientific ComputingScience Research

Inherent Lab Introduces Faraday, an AI Agent That Reproduces Computational Experiments from Research Papers

16 August 2026· 260816011

Inherent Lab Introduces Faraday, an AI Agent That Reproduces Computational Experiments from Research Papers

On August 14, Inherent introduced Faraday, a 27 billion parameter model fine-tuned on the Replica task set. The agent reads a paper with one results figure removed, then has one hour to write and run code intended to reproduce the experiment. Across 68 held-out AI for science tasks, Faraday received a higher score under the authors’ evaluation framework than the other two models, Claude Opus 4.8 and GPT-5.5 Codex, in 60% of the tasks.

“Scientific papers describe what worked for their authors, not the failed results that led them there,” Inherent writes in its research announcement.

A published paper usually presents the successful method and the final figure. Reconstructing the work requires retracing the hypotheses, trial runs, and simplifications that still test the original idea.

In a July experiment using a chain of language models in computational physics, the system first reproduced published calculations and checked its numbers against the literature, then ran its own calculations. Replica turns this type of verification into a task for training an agent.

In the preprint, the authors selected 100 published papers on machine learning and AI for science and created 310 tasks. For each task, they remove one results figure from the PDF. The agent reads the remaining text, receives access to a computational environment and a coding assistant, then constructs a minimal version of the experiment that tests the paper’s claim again.

Before the run, Claude Opus 4.7 prepares a rubric for each task, using only the text of the paper. GPT-5.5 Codex then examines the code, action log, working environment, and original figure. It reruns the code when necessary. Figure similarity is one of five criteria. The other criteria assess whether the experiment supports the scientific claim, reproduces the design of the original experiment, uses the computational budget appropriately, and follows the task rules.

Here, the completed plot is used to evaluate how the agent obtained the result.

In one analysis, Faraday searched a mathematical map generated by a molecular model for regions from which the model could construct new structures. The task therefore tested whether the model could be used to discover new molecules. Selecting existing molecules from a dataset would only demonstrate the ability to retrieve already known candidates.

Faraday was trained using reinforcement learning. The mean score from three evaluations guided the model’s next training update, while the algorithm associated that score with the decisions the agent made during the one-hour run. The authors gave Faraday and the two competing models the same papers, one hour, and the same computational budget. Each agent was run eight times on every task.

The judge receives the code and action log together with the original figure, which allows it to examine the experimental process that produced the result.

Originally published on Telegram by Ukhvat NewsView on Telegram
Sources
#faraday#inherent-lab#replica-benchmark#experiment-reproduction#reinforcement-learning#ai-agents