Live·Open questions in longevity research
All news
Scientific ComputingScience Research

Benchling tested whether AI can revise laboratory protocols. The best score was 59.2%

16 August 2026· 260816008

Benchling tested whether AI can revise laboratory protocols. The best score was 59.2%

On 13 August, Benchling released BenchBench-Protocol, a benchmark based on changes made to 96 published laboratory protocols. The model must propose a revision for a specific experiment and account for its effects on subsequent steps. Nine models were tested. Claude Opus 5 achieved the highest average score, receiving 59.2% of the maximum under the authors’ weighted scoring system.

A published protocol rarely meets a laboratory’s needs without modification. A researcher may use a different instrument, substitute a reagent, or observe an unexpected result. The researcher must then adjust the next step so that an earlier problem does not compromise the final measurement.

“Language models know a surprising amount about biology. But knowing biology and doing biology are different things,” writes Benchling cofounder and president Ashu Singhal.

The BenchBench authors compared published protocols with versions that scientists had modified during actual laboratory work. For each difference, they created a question for the model and a set of evaluation criteria. Each criterion was weighted according to the specific revision. At least two independent specialists in the relevant field reviewed every task, and each reviewer had at least three years of laboratory experience.

For each task, the model receives a request written from a researcher’s perspective, the original protocol, and internet access. It then provides a free-form response. In Benchling’s examples, models omitted an instrument setting, treated a contaminated sample as clean, or failed to recognize a change in technique that could cause the experiment to fail.

Claude Opus 5 achieved an average score of 59.2%. The other systems scored between 34.1% and 47.1%. Each model’s score was averaged across ten attempts on all 149 tasks. The result measures how well a model can propose a revision that accounts for the experimental conditions and the effects on subsequent steps.

Originally published on Telegram by Ukhvat NewsView on Telegram
Sources
#benchling#benchbench-protocol#laboratory-protocols#claude-opus-5#wet-lab-ai#protocol-revision