Live·Open questions in longevity research
All news
AI in medicineLongevity research

An anonymous research team, some members claiming Stanford affiliation, has launched RejuvenationBench, an AI benchmark whose final stage tests proposals on living cells: GPT-6 Astra and Muse Spark 1.3 lead in scientific evidence evaluation, while Claude Opus 5.5 refused all 24 such tasks despite nearly matching the leaders in calculations and experiment planning

29 September 2026· 260929007

An anonymous research team, some members claiming Stanford affiliation, has launched RejuvenationBench, an AI benchmark whose final stage tests proposals on living cells: GPT-6 Astra and Muse Spark 1.3 lead in scientific evidence evaluation, while Claude Opus 5.5 refused all 24 such tasks despite nearly matching the leaders in calculations and experiment planning

The project describes itself as the first AI benchmark that extends model evaluation all the way to laboratory testing on living cells: nine systems have already completed 96 tasks covering evidence comprehension, effect calculation, experiment design, and outcome prediction. Starting in early 2027, the two top-performing models will propose an intervention for laboratory verification under biosafety oversight.

Previous benchmarks in geroscience mostly test whether an AI can calculate biological age from lab results, as with LongevityBench, which compared models on predicting age from blood panels and DNA methylation. The team behind RejuvenationBench, volunteers and researchers (the site explicitly states that some team members' connection to Stanford University does not imply the university's endorsement or sponsorship), set out to test something different: whether an AI assistant can conduct genuine rejuvenation research from initial data to a final conclusion. The benchmark is divided into four sections: understanding what an experiment actually proves and not confusing cells in a dish with a human subject; correctly calculating the effect of an intervention while accounting for cell donors and missing data; identifying hidden confounders and choosing an experiment that genuinely reduces uncertainty; and providing a quantitative prediction of the outcome.

The version 0.2 snapshot from September 28 is still a calibration release, but it already revealed a sharp spread among nine models from OpenAI, Anthropic, Meta, xAI, DeepSeek, and Moonshot AI. In evidence evaluation, GPT-6 Astra and Muse Spark 1.3 (Meta's model) lead, both solving 21 out of 24 tasks in full. GPT-6 Luna and DeepSeek V4 Pro follow, while GPT-5 Nano barely manages. Claude Opus 5.5 scored exactly zero in this section: the model refused to answer all 24 tasks, even though the same model placed second in effect calculation and third in experiment design, nearly matching the leaders. This appears to be Anthropic's safety filters at work. The same company has previously blocked questions in Fable 5 such as ten unsolved problems in cancer research, treating them as a biosafety threat: the benchmark authors are discussing with Anthropic how to reduce refusal rates without sacrificing caution. For comparison, Claude Haiku 4.5 answered every question without refusing but often gave wrong answers.

The team does not publish the full set of correct answers, so that models cannot be trained to memorize the tasks instead of learning to solve them: a written test is vulnerable to rote learning, while a live laboratory experiment tests genuine capability. A pilot round already took place in 2026: models were limited to FDA-approved compounds only, as a check on how safe their choices were before allowing free selection. The round in which the models themselves choose the experimental candidate, which the authors call Livematch, is planned for 2027: only the two top-performing models from the current calibration stage will be admitted, and scientists will review the proposed experiment for safety before it goes to the laboratory.

The authors explicitly note that high scores reflect a model's research reasoning, and the definitive answer about the real-world effect of an intervention will come only from the live round in 2027.

Originally published on Telegram by Ukhvat NewsView on Telegram
Sources
#rejuvenationbench#ai-benchmark#gpt-6-astra#claude-opus-5#geroscience#rejuvenation-research