Mol-JEPA Predicts Drug Molecule Properties from Molecular Structure and Biological Data
Mol-JEPA Predicts Drug Molecule Properties from Molecular Structure and Biological Data
On August 23, the authors published a preprint describing Mol-JEPA. They compiled 14 types of data for 4.69 million small molecules, ranging from molecular structures to cellular measurements and ADMET profiles, which describe absorption, distribution, metabolism, excretion, and toxicity. During training, the model randomly hides one type of data and uses the remaining types to reconstruct its compressed numerical representation.
A drug candidate may bind well to its intended protein but leave the body too quickly or prove toxic. A small structural modification can sometimes cause a large change in one of these properties. The authors therefore preserve each molecular structure without distortion and train the model to connect different observations of the same molecule.
Mol-JEPA receives atom and bond diagrams, calculated chemical features, biological assay results, cellular profiles, and ADMET data. Separate components convert each type of data into a short numerical representation with a common format. A shared module then receives the available representations and predicts the representation of the hidden data type. Through this process, the model learns to connect molecular structure with biological and pharmacological properties.
In property prediction tasks involving some small datasets, Mol-JEPA produced a lower mean absolute error than the comparison methods. In a separate evaluation, its advantage increased as the structural difference between the test molecules and the training portion of the task grew. On public temporal splits, TabICLv2, a model for tabular data, outperformed Mol-JEPA more often, both in mean absolute error and in pairwise comparisons.
With a random split, molecules that share a structural core may appear in both the training and evaluation sets, which can make performance estimates overly optimistic. The catalog of evaluation methods for drug development specifically warns about this risk in small, noisy datasets. Evaluating Mol-JEPA on structurally distant compounds therefore provides a more accurate test of whether its advantage extends to molecules outside the training portion of the task.
A separate experiment showed that additional information affects prediction accuracy. The authors retrained one version using the molecular graph and ECFP4, a digital fingerprint of molecular structure, and another version using all available data types. In subsequent tests, they used each representation as input to two prediction models. The full version achieved a mean absolute error that was 14% lower with the simpler model and 13% lower with the more complex model.
Drug development requires connecting molecular structure with accumulated evidence about how a molecule behaves. Mol-JEPA combines these data into a single numerical representation that can then be used to predict the properties of drug candidates.