Live·Open questions in longevity research
All news
AI in medicineScience Research

A model trained on antibody target contact regions predicted binding strength more accurately

18 August 2026· 260818013

A model trained on antibody target contact regions predicted binding strength more accurately

On August 13, Communications AI & Computing published a study of a language model that reads the sequences of both antibody chains. During training, the authors masked amino acids primarily in the six CDR loops, where the antibody contacts its target, and tested the predictions on antibody variants against six antigens. In a dataset of 11 052 variants, predictive performance improved by 26,6% relative to the original model.

An antibody recognizes its target through the ends of two protein chains. Each end contains three CDR loops, which together form the contact surface. The remaining parts of the chains hold these loops in the required shape. Replacing an amino acid in a loop can change binding strength even when the rest of the sequence remains almost unchanged.

A protein language model learns to reconstruct masked amino acids in a sequence. In conventional training, masked positions are distributed randomly across the entire chain. As a result, some training tasks involve framework regions, which are relatively similar across antibodies. CDR loops are more diverse and determine which target an antibody recognizes. The Boston University team trained the model on paired heavy and light chains and masked 50% of the amino acids within the CDR loops. The resulting number of masked positions matched the conventional masking of 15% of the entire chain.

The training dataset contained 1,6 million naturally paired human antibody sequences from the Observed Antibody Space database. After training, the model converted each chain pair into a set of numbers. Linear regression applied to this representation estimated the logarithm of the dissociation constant KD. A lower KD indicates stronger binding between the antibody and its target. This approach allowed the authors to test whether the paired-chain representation could rank antibody variants with existing measurements according to binding strength.

In the dataset of antibody variants against fluorescein, the proportion of variation in the measurements explained by the prediction increased from 0,547 to 0,693. In a dataset of 71 830 antibody variants against a region of the SARS-CoV-2 S protein, this measure increased from 0,366 to 0,396, while the mean absolute deviation between predictions and measurements decreased from 0,865 to 0,841.

This training strategy follows from antibody structure. The original model reconstructed framework regions with an accuracy of 72–92%, but reconstructed the most diverse CDR loop correctly in only 35,69% of cases. The authors attribute the improvement to CDR masking, which teaches the model patterns of amino acid combinations within the loops and dependencies between the two chains.

The authors also tested whether training ESM2 on a very large dataset of individual chains would improve the language model. They first trained it on 1,2 billion individual-chain sequences and then on paired chains. After the second stage, its results converged with those of the version trained on paired chains from the outset. Across several datasets, the ESM C language model, which has five times fewer parameters, performed comparably to larger specialized models.

Originally published on Telegram by Ukhvat NewsView on Telegram
Sources
#antibody-language-model#cdr-loops#binding-affinity#protein-language-model#dissociation-constant#antibody-variants