Live·Open questions in longevity research
All news
AI in medicineScientific ComputingLongevity research

Insilico launches DDD Benchmark, a service for comparing AI models on drug development tasks

31 July 2026· 260810117

Insilico launches DDD Benchmark, a service for comparing AI models on drug development tasks

On July 30, Insilico Medicine launched DDD Benchmark. Developers can submit a model and receive a private report or publish its results in a public leaderboard.

A single prediction cannot identify a drug. Researchers first select a protein or cellular process associated with a disease. Chemists then search for a molecule and test its properties and its ability to reach the target tissue. The team must subsequently decide whom to enroll in a clinical trial and which outcome to measure. At each stage, models receive different data and address different tasks.

DDD Benchmark brings these tasks together in one service. As of July 31, the public catalog contains 243 tasks across six categories: disease biology, medicinal and synthetic chemistry, clinical trials, longevity, and therapeutic biomolecules. The tasks include predicting molecular properties, selecting a biological target, reconstructing the sequence of reactions needed to synthesize a compound, and predicting the outcome of a clinical trial. Instead of assigning a general score to a “model for pharma,” the service reports performance at individual stages of drug development.

In June, Insilico evaluated GPT 5.5 on 162 internal tasks. The new service accepts models from external developers and compares their results in a shared leaderboard. In its announcement, Insilico says it will accept any model that operates through a standard API interface.

The longevity category tests how models work with specific biomedical data. In one task, a model receives DNA methylation levels from age-associated genomic regions in two blood donors and determines which donor is older. Other tasks use gene activity, NHANES biomarkers, blood proteins, and SynergyAge lifespan data. This gives aging tasks their own entries in model comparisons instead of placing them in a general dataset without context.

The public leaderboard will show which models perform best at target selection, molecular property prediction, synthesis, or clinical outcome prediction. This gives developers a common set of tasks instead of relying on company demonstrations that cannot be compared directly.

Originally published on Telegram by Ukhvat NewsView on Telegram
Sources
#drug-discovery#ai-benchmark#model-leaderboard#target-selection#molecular-properties#longevity-data