Insilico Medicine launches O3DC consortium with an open catalog of 115 benchmarks for AI in drug development
Insilico Medicine launches O3DC consortium with an open catalog of 115 benchmarks for AI in drug development
On August 24, Insilico Medicine announced the launch of the Open Drug Discovery & Development Consortium (O3DC), an open consortium for evaluating AI in drug development. Its first resource, the Benchmark Index, initially brought together 115 open benchmarks across ten categories, along with 15 consortia and initiatives.
A model’s evaluation is meaningful only when the testing conditions are known, including the task assigned to the model and the method used to compare its answers. A benchmark is a set of such tasks and evaluation rules.
“AI evaluation for drug discovery is scattered across dozens of repositories, competition pages, and leaderboards that have not been updated in years,” Insilico Medicine writes.
In July, Insilico launched the DDD Benchmark, which compares models across individual stages of drug development. This leaderboard shows the score a model received on a specific task. The Benchmark Index describes the conditions under which that score was obtained, including what the benchmark measures and which methodological feature may make the score a poor guide to performance.
In the Benchmark Index, each benchmark is listed with its task, maintainer, code, repository status, and a known limitation that may distort the result. This information makes it possible to connect each score with the rules used to produce it.
One of these rules concerns how the data are split. The split determines which examples the model learns from and which are reserved for evaluation. The Therapeutics Data Commons (TDC), a resource that provides tasks, curated datasets, and benchmarks for drug development, uses predefined, methodologically appropriate splits.
In the MoleculeNet entry, which describes a machine learning benchmark for molecules, O3DC warns about small datasets and noisy labels. With a random split, molecules that share a chemical substructure may appear in both the training data and the evaluation set. The evaluation may then overestimate the model’s ability to work with new molecules. The entry therefore recommends splitting the data according to the molecules’ shared structural scaffold.
Before relying on a model’s score, readers can now check which task the model was given, how the data were split, and where the benchmark itself may distort the result.