Live·Open questions in longevity research
All news
AI in medicineScience Research

Across two single-cell datasets, four AI models predicted surface proteins from RNA more accurately than highly variable gene selection

8 August 2026· 260810029

Across two single-cell datasets, four AI models predicted surface proteins from RNA more accurately than highly variable gene selection

On August 7, the authors of a bioRxiv preprint compared scGPT, SCimilarity, UCE, and Transcriptformer across nine public single-cell datasets. They kept the model weights fixed and tested whether the models’ numerical cell representations performed better on new tasks than highly variable gene selection. All four models showed an advantage in predicting surface proteins.

Each cell has an RNA profile, which records how actively each gene is transcribed in that cell. In a bioRxiv preprint, the authors tested whether numerical representations of these profiles performed better on new tasks than highly variable gene selection.

The AI model weights remained fixed. Separate simple algorithms were trained on the numerical cell representations to identify cell types or predict protein levels. When combining data from different experiments, the researchers compared these representations with specialized software. This design separated knowledge acquired during the models’ earlier training from learning specific to each new task.

In two datasets that measured both RNA and surface proteins in each cell, all four models produced more accurate predictions than highly variable gene selection. First, each model converted an RNA profile into a set of numbers. Linear regression, a simple prediction algorithm, was then trained on paired RNA and protein measurements and used that numerical representation to predict protein levels. Correlation measures how closely the predictions track variation in protein levels, while error measures how far they are from the observed values.

In umbilical cord blood cells, scGPT achieved a median correlation, meaning the central value among the resulting estimates, of 0.760 between predicted and measured values, compared with 0.612 for the conventional method. The error was about 0.32 for scGPT and SCimilarity, compared with 0.438. In bone marrow cells, Transcriptformer achieved a correlation of 0.768, compared with 0.575 for the conventional method, while SCimilarity reduced the error from 0.536 to 0.181.

For cell-type classification, highly variable gene selection often produced equally accurate results. When combining data from different experiments, scVI, a program designed for this task, was better at integrating the measurements into a common representation, while SCimilarity preserved differences between cell types more effectively in several datasets. As the authors write, “Their effectiveness depends on the task and evaluation conditions.”

Originally published on Telegram by Ukhvat NewsView on Telegram
Sources
#single-cell-rna#surface-proteins#scgpt#scimilarity#transcriptformer#foundation-models