Shift Bioscience shows that virtual cell AI models work, attributing their apparent failure in 2025 to a poor evaluation metric
Shift Bioscience shows that virtual cell AI models work, attributing their apparent failure in 2025 to a poor evaluation metric
On October 1, Nature Biotechnology published a paper by UK company Shift Bioscience and University of Toronto professor Bo Wang. It attributes last year's conclusion that AI models failed to predict cellular responses to gene silencing to poorly calibrated evaluation metrics. With a calibrated metric, most models outperformed baseline predictions. Shift Bioscience will use the metric to identify targets for interventions against aging and fibrosis, the excessive growth of scar tissue.
A “virtual cell” is an AI model that predicts how gene activity changes when a gene is switched off or on, allowing researchers to screen out unpromising targets computationally. In 2025, several groups, including the authors of a paper in Nature Methods, concluded that neural networks such as scGPT and GEARS often failed to outperform the simplest prediction: using average gene activity from previous experiments as the predicted response to any new gene perturbation. The same team's July analysis had already identified the evaluation metric itself as the main suspect.
The team found an explanation in a failure of its earlier method. A technical replicate, which uses half the cells subjected to an intervention to predict the other half, necessarily contains a signal and should outperform a null prediction that assumes “nothing changed.” Yet on the Replogle22 K562 dataset, it scored worse than the null prediction in 95% of cases under the conventional error metric. In that dataset, silencing a gene significantly changes the activity of only 3,25 genes on average out of the 20 thousand active genes that make up the cell's transcriptome, compared with 102 in Norman19. The conventional metric gives every gene equal weight, so noise overwhelms the sparse signal. The null prediction receives a misleadingly good score, while an accurate prediction scores worse.
To measure this weakness, the authors constructed an accurate positive control called an interpolated replicate. It combines the technical replicate with the mean prediction in proportion to the strength of the perturbation. They also introduced dynamic range fraction (DRF), a metric that compares this control with the null prediction under a particular error metric. A high DRF means the error metric reliably distinguishes a good prediction from a meaningless one; a low DRF means the two receive almost identical scores.
Across 14 datasets, conventional metrics were least able to distinguish prediction quality when a perturbation changed only a few genes. Calibrated metrics gave a different picture: most of the nine models evaluated, including scGPT and GEARS, outperformed baseline predictions, while newer architectures, PRESAGE, scLambda and CellFlow, did so almost always. The result held on an independent dataset, and high scores coincided with recovery of biological pathways, sets of reactions that a cell coordinates to perform a particular task.
“These models are more than toys. The field just needed better rulers,” Bo Wang wrote on X.
Since 2017, Shift Bioscience has been seeking genes that can rejuvenate cells without the risk of tumors and working to combat fibrosis. According to chief scientific officer Brendan Swain, the calibrated metric increases confidence in the models' predictions. The company has already identified its first dual-purpose target, the gene SB-101, suitable for both rejuvenation and the treatment of fibrosis, and is now searching for others with the same potential. Stanford associate professor Anshul Kundaje described the work as rigorous, while noting that it should have been done sooner: better late than never. The authors identify transferring a model to new cell types as an unresolved challenge.