Live·Open questions in longevity research
All news
Scientific ComputingScience Research

A compressed ESM-2 model retained its average accuracy but severely distorted the ranking of substitutions in UBR5

18 August 2026· 260818005

A compressed ESM-2 model retained its average accuracy but severely distorted the ranking of substitutions in UBR5

On August 15, Qin Shao published a preprint comparing six computational modes across three versions of ESM-2. The evaluation covered 201 sets of laboratory measurements from ProteinGym, a dataset of the effects of amino acid substitutions in proteins.

ESM-2 reads an amino acid sequence and ranks possible substitutions by their expected effects. A laboratory may select the highest-ranked candidates for its next experiment, so it needs the variants for its target protein to be ordered accurately.

Shao processed 2,41 million variants using three ESM-2 model sizes and six precision modes. In the 3 billion parameter model, 8-bit compression had almost no effect on the average correlation with laboratory measurements: the difference from full precision was −0,0021. For the human protein UBR5, however, the same correlation fell from 0,591 to 0,223. The average metric described performance across the entire set of experiments, while the ranking of variants for one protein changed substantially.

The reason lies in how the score is calculated. The model scores a substitution by subtracting the similar log probabilities of the original and new amino acids. After subtraction, a small compression error in each value can account for a substantial part of the final difference and change the order of the candidates.

The author also replaced symmetric scaling with asymmetric scaling, which changes how the 8-bit format distributes the model's internal values. In this evaluation, the largest deviation from full precision decreased from −0,3688 to −0,0395, and it did not exceed 0,05 in any of the 201 experiments.

Shao proposes running the variants of the target protein through both the full-precision and compressed versions of ESM-2 before the next round of laboratory experiments, then comparing the two rankings. This screening procedure uses the output of the full model and requires no new laboratory measurements.

Of the 18 combinations of model size and precision, only three configurations of the 650 million parameter ESM-2 model reached the Pareto frontier, the set of options that cannot be simultaneously surpassed in average accuracy, memory use, and speed. For this task, model size determines the balance between average accuracy and computational resources, while the compression mode determines the risk of distorting the ranking for a specific protein.

Originally published on Telegram by Ukhvat NewsView on Telegram
Sources
#esm-2#protein-gym#model-quantization#protein-variant-ranking#ubr5#asymmetric-scaling