Live·Open questions in longevity research
All news
Scientific ComputingScience Research

Biodyn: 109 problems for interpreting biology inside models

15 August 2026· 260815009

Biodyn has published 109 problems that ask how laboratory experiments can validate signals found inside AI models of biology

On August 14, BiodynAI program lead Igor Kendyukhov published a map of 109 open problems. It defines six tests that a feature identified inside a model should pass before researchers use it to select a laboratory experiment.

These AI models are trained on DNA and protein sequences, as well as single-cell data. A researcher may find a feature in the model's internal computations that resembles a known relationship between genes or a known cell state. Agreement with a reference map shows that the researcher has extracted a familiar pattern from the model. Further tests are needed before that pattern can support a biological explanation.

The Biodyn map divides this process into six stages. First, researchers extract a biological property from the model and give it a meaningful interpretation. They then test whether the feature provides information beyond simple explanations and intervene on the corresponding model component to determine its role in the computation. Next, they reproduce the result using other data, models, or species. At the final stage, the model must predict a previously unknown biological fact in advance, and a laboratory experiment must confirm it.

The map's logic grew out of Kendyukhov's July paper, which tested a model's attention mechanism in two tasks. Separate analyses by cell state helped recover known regulatory relationships recorded in reference databases. In the task of predicting which genes would change activity after CRISPRi or CRISPRa, methods that respectively suppress or increase gene activity, attention weights did not improve predictions beyond simple gene properties, such as mean activity.

An independent review distinguishes three levels. Information can first be read from a model's internal state. Researchers can then determine the role of a model component in the computation. A causal mechanistic explanation requires experimental validation.

At the sixth stage, the model selects in advance the experiment expected to reduce uncertainty most: which gene knockout, drug dose, or measurement timepoint to test. Researchers record the prediction before obtaining the result. The laboratory experiment then shows whether the model's internal signal led to new biological knowledge.

Sources
#mechanistic-interpretability#biological-ai#foundation-models#experimental-validation#crispri#single-cell