Live·Open questions in longevity research
All news
AI in medicineScientific ComputingScience Research

AI models predict post-knockout gene expression accurately but fail at cell-type composition: Pop-Corn, a new method validated in cultured cells and living tissue, solves this problem

29 September 2026· 260929001

AI models predict post-knockout gene expression accurately but fail at cell-type composition: Pop-Corn, a new method validated in cultured cells and living tissue, solves this problem

On September 26, researchers from Yale University and the Whitehead Institute for Biomedical Research published a preprint describing the Pop-Corn method. In a benchmark on nearly 58,000 human T cells, the method recovered cell-type and cell-state composition after perturbation more accurately than four alternatives and better preserved rare states. It was also separately validated on sections of living mouse liver tissue.

Perturb-seq is an experiment in which different genes are knocked out simultaneously across a pool of cells and the expression of the remaining genes is read out in each individual cell. Testing every gene is infeasible, so a model is needed that can predict the effect of an untested perturbation and guide what to test next.

The leading models, GEARS, scGPT, CPA, and CellFlow, approach the problem in two steps: they predict gene expression for a cell, then derive cell-type and cell-state composition from that prediction after the fact.

On human T cells, this logic broke down. All four predicted mean gene expression reasonably well, yet when the predictions were converted into composition, the proportions were wrong: the models assigned cells to the most common state and missed rare ones. Correlation of predicted composition with the ground truth: CPA 0.187, scGPT 0.629, CellFlow 0.681, GEARS 0.709, Pop-Corn 0.802, with lower error and a closer match to reality.

The practical utility of the method was tested with the very task it is designed for: genes were ranked by their predicted effect on cell state, and the top 5 were compared with the actual top 5. The predictions matched in 70 out of 310 cases, 22.6%.

Expression models fail here for a straightforward reason: they predict an average gene-expression profile, and when that profile is translated into cell-type composition by matching to the nearest similar cells, rare states are lost because they do not resemble the average. A similar weakness of GEARS and scGPT was already shown in July by the Needles in the Haystack analysis: on gene expression alone, a naive "dataset-mean profile" baseline outperformed them, because the error metric weights all genes equally even though a perturbation changes only a subset. Pop-Corn skips this step entirely: it treats all cells sharing one perturbation as a single sample and, instead of encoding the target by gene ID, encodes it by the protein sequence of its gene product using the ESM-2 model (Meta AI). This also works across species, because human and mouse genes encode similar proteins.

This cross-species transferability was tested on sections of living mouse liver, where 202 genes were knocked out. In intact tissue, a cell's response depends not only on its own knocked-out gene but also on neighboring cells, something invisible when cells are dissociated. By adding an encoding of each cell's position within the tissue section, the authors predicted neighborhood composition with a correlation of 0.93 and the change in composition caused by the perturbation with a correlation of 0.77.

The model's internal attention (which neighboring cells receive the most weight) pointed to five genes with the strongest effect on surrounding cell types. For Dnm2, the model highlighted endothelial cells lining the vasculature, consistent with the gene's known role in blood vessel formation. For Snip1, it highlighted hepatocytes, though this has not yet been confirmed. For the remaining three genes, no selective effect was found. The authors describe these as hypotheses for future validation.

Pop-Corn reveals what changed in the composition of a cell population. Identifying which gene or pathway triggers that change is a question for expression models, and the authors name the integration of the two approaches as the next step.

Originally published on Telegram by Ukhvat NewsView on Telegram
Sources
#perturb-seq#gene-knockout#cell-composition#esm-2#single-cell-rna#pop-corn