SciLaws-Bench separates formula fitting from scientific law discovery
Across 3,616 pairwise comparisons in SciLaws-Bench, a more accurate fit corresponded to greater scientific validity in only 54.9% of cases.
In 3,616 pairwise comparisons, a more accurate fit corresponded to greater scientific validity in only 54.9% of cases. The authors obtained this result while evaluating nine models with SciLaws-Bench, a preprint published on September 1. SciLaws-Bench contains 118 tasks drawn from 381 scientific papers. It evaluates predictive accuracy and compliance with the requirements of the studied phenomenon separately.
A nuclear physics task shows how a model can fail. The most accurate formula introduced a resonance peak absent from the original phenomenon. It predicted held-out data more accurately while violating a known property of the phenomenon.
The benchmark has two settings. In SciLaws-Real, a model receives real observations and proposes a formula. The formula is evaluated on held-out data and checked for the required sign of a quantity, known limits, and mandatory dependencies. A language-model judge evaluates these criteria. The authors compared its decisions with assessments from five domain experts.
In SciLaws-Parallel, the authors create a synthetic world governed by a new version of a published formula. The model selects its own measurement points, receives noisy responses, and reconstructs the hidden law. This setting also tests whether the model can choose measurements whose results distinguish between competing formulas.
After evaluating nine models, the authors write: “Models are better at generating laws than at selecting them reliably.” The work is published as a preprint, and a language-model judge evaluates part of the criteria. The result therefore identifies a specific gap between predictive fit and scientific validity within the SciLaws-Bench tasks.
The page for the organization Experiment: Experiment on Eternal Search.
Open the related Eternal Search page
Sources
[1] yiyihum.github.io
SciLaws-Bench introduces a concrete, quantified gap between formula-fitting accuracy and scientific validity across 3,616 comparisons, a distinction directly relevant to evaluating AI in research contexts; the Experiment organization page in Eternal Search is connected to this work and gives readers a path from a specific benchmark number to an organizational profile they can inspect.