Talus Bioscience presents Ptarmigan-1: a model that finds molecules for proteins without predicting their three-dimensional binding poses
Talus Bioscience presents Ptarmigan-1: a model that finds molecules for proteins without predicting their three-dimensional binding poses
On July 30, Talus Bioscience published a preprint on Ptarmigan-1 on bioRxiv. The model matches a protein's amino acid sequence to a molecule's two-dimensional chemical representation and identifies the protein residues that it predicts will interact with the molecule.
Drug discovery begins with a library of molecules and a simple question: which of them should be synthesized and tested in the laboratory? A computer would usually predict a three-dimensional pose for each candidate, showing where the molecule would bind within the protein and which atoms it would contact. This calculation provides a detailed analysis of a known pocket, but applying it to billions of candidates requires substantial computing resources.
Ptarmigan-1 first screens candidates using the protein sequence and the molecule's two-dimensional formula. The model converts each amino acid residue and each molecule into a set of numbers. The proximity between these sets determines the interaction score. It directly produces a ranked list of molecules and a map of the residues that it predicts will contact each candidate.
The main technique is to encode the entire molecular library in advance and store it as an index. When a new protein sequence becomes available, the model only needs to search the index for molecules that are close to its residues. The authors report that they screened 3.4 billion compounds from OnePot CORE against 20,431 proteins in the human proteome in less than one day, using 20 NVIDIA H100 GPU-hours.
This approach changes the order of operations: rapid screening covers a very large library, while detailed analysis is reserved for a short list. Researchers can then predict three-dimensional poses for the selected pairs, synthesize the molecules, and test them experimentally.
This sequence is particularly useful for difficult targets. A cryptic pocket opens only when it comes into contact with a molecule. A disordered protein region changes shape continuously and often does not form a stable cavity. The authors also trained Ptarmigan-1 on measurements that establish either the interaction itself or the identity of the modified residue. The model can therefore handle such cases without a previously determined three-dimensional structure.
In the authors' table, Ptarmigan-1 achieved an adjusted logAUC of 0.12 across five structured LIT-PCBA targets, compared with 0.196 for Boltz-2. On a covalent dataset containing seven disordered targets, the mean ROC-AUC was 0.68 for Ptarmigan-1 and 0.44 for Boltz-2. The model is useful as a broad filter for finding new pairs of proteins and molecules, including interactions involving regions for which structure-based calculations have no suitable starting point.