Live·Open questions in longevity research
All news
AI in medicineTherapeutics

Gero founder Peter Fedichev says his model designed inhibitors of two kinases using only protein amino acid sequences, without three-dimensional structures, and confirmed their activity in vitro

2 October 2026· 261002005

Gero founder Peter Fedichev says his model designed inhibitors of two kinases using only protein amino acid sequences, without three-dimensional structures, and confirmed their activity in vitro

Gero is Peter Fedichev’s biotechnology company, which uses AI models of aging to identify drug targets. On October 1, he wrote that two papers from his team had been accepted at a NeurIPS 2026 workshop in Sydney on AI for drug discovery. He said the new version of ProtoBind-Diff had taken the method through to purchased molecules tested in vitro for the first time, yielding inhibitors of two kinases, EGFR and TYK2.

Most drug design models predict how a molecule will fit into a binding pocket on a protein’s surface, which requires the protein’s three-dimensional structure. Fewer than 30 thousand structures are available for protein-drug pairs, and almost all cover a small number of proteins. Entire classes, such as intrinsically disordered proteins that lack a single stable shape, have no suitable structure at all. Fedichev describes this shortage as follows:

I think this is the most revealing bottleneck in the field because almost everyone treats it as a law of nature, when it reflects a choice.

ProtoBind-Diff works without a structure: it takes a protein sequence as input and produces a candidate molecule. This is possible because the sequence encodes almost all the necessary information about the protein, as demonstrated by the language model ESM-2, whose protein embedding ProtoBind-Diff uses. That embedding feeds into a diffusion network. The principle is the same as in image generators such as Midjourney, except that the network builds a chemical formula from noise instead of pixels. In the preprint, the team acknowledges that its approach using graph models trained on structural data did not work: with too few examples, the model tended to generate meaninglessly long molecules. Switching to a text representation resolved this problem. Without access to a 3D structure, the model’s attention heads learned to identify amino acids in the binding pocket, achieving a ROC-AUC of 0,72 on a scale where 0,5 represents random guessing and 1 represents perfect prediction.

The same model weights also provide a screening tool without further training: a function that uses the protein sequence to estimate whether an existing molecule will bind to it. According to Fedichev, it distinguished active from inactive compounds more accurately than docking, with a ROC-AUC of 0,70 compared with 0,54 for Vina and 0,60 for Boltz-1. This comparison is not yet included in the published preprint.

Fedichev says the main new result for NeurIPS is that the team used this scoring function to screen a catalogue of commercially available compounds, then purchased and tested a subset, finding inhibitors of two kinases, enzymes that activate other proteins through phosphorylation. The best compounds inhibited EGFR, a target of lung cancer drugs, at 0,44 micromolar and TYK2, a target of psoriasis drugs, at 0,75 micromolar. Their chemical scaffolds had not previously been reported as inhibitors of any kinase.

OmniSyn had already tried the same approach in September, producing a library of candidate molecules for 21 thousand human proteins, but the work remained at the computational evaluation stage. Fedichev claims his team was the first to reach the point of purchasing compounds and measuring their activity in vitro.

The data ProtoBind-Diff needed were buried in patents that almost no one analyzes. To extract them, the team launched the HARVEST multi-agent pipeline in March. It reads binding data tables directly from US patent office archives.

Data were the bottleneck. Structure had never been the limiting factor.

He says a private version of ProtoBind-Diff trained on the full HARVEST dataset is already showing signs of generalizing to proteins it did not encounter during training. This is harder than selecting molecules for a familiar target. The team plans to use this version for aging-related targets.

Originally published on Telegram by Ukhvat NewsView on Telegram
Sources
#protobind-diff#drug-design#egfr#tyk2#kinase-inhibitors#in-vitro