Live·Open questions in longevity research
All news
Scientific ComputingScience Research

Paper Proposes Evaluating AI Scientists Through Their Record of Research Decisions

22 August 2026· 260822013

Paper Proposes Evaluating AI Scientists Through Their Record of Research Decisions

On August 18, nine authors posted a proposed framework on ChemRxiv, a preprint repository, for evaluating scientific AI agents, which are programs that select their own research steps. The framework uses the “discovery episode” as its unit of evaluation: a recorded sequence of hypotheses, actions, data, and revisions.

A conventional exam problem specifies a question, an expected answer format, and a scoring scale. In research, an AI agent formulates a hypothesis, selects a calculation or experiment, obtains data, and decides which step could advance the investigation. The authors propose evaluating this entire sequence.

In an autonomous research benchmark, agents changed the training settings of a language model, ran calculations, rechecked ambiguous results, and returned to earlier approaches that had failed. The authors propose treating this sequence of decisions, data, and revisions as the unit of evaluation.

The authors call this record a discovery episode. It captures what was known before each new step, which action the agent selected, what it observed, and how its plan changed afterward. Each step records the code version, instrument settings, human involvement, and safety rules. Another researcher can use this record to reconstruct the conditions under which the data were generated, examine how the work proceeded, and repeat the calculation or experiment.

The framework divides research into three connected parts. A hypothesis must lead to a testable outcome. Execution shows whether the agent can perform a calculation or experiment with the available resources and recover from a failure. Interpretation connects the data to the hypothesis and determines the next step when the results conflict or remain ambiguous.

An episode includes null and anomalous results, as well as execution failures. It preserves cases in which a test produced a null or unusual result and records what the agent did after an error. The authors propose storing these records alongside successful steps so that scientific agents can be trained and evaluated using the complete trajectory.

The authors recommend starting with bounded tasks that have a clear objective, an outcome that can be scored automatically, and a calculation or experiment that can be completed within acceptable limits of time and cost. Individual stages can then be combined into episodes, and the outcome can be verified through independent replication. This allows an evaluator to trace the path from a decision to the resulting data and then to the next research question.

Originally published on Telegram by Ukhvat NewsView on Telegram
Sources
#ai-agents#scientific-discovery#research-evaluation#discovery-episodes#autonomous-research