LifeSciBench: best AI passes 36% of real biology tasks
LifeSciBench tests AI on 750 real biology tasks — top score is 36.1%, 171 tasks stumped every model, and performance drops further on anything beyond plain text.
GPT-Rosalind scored highest in LifeSciBench — passed 36.1% of tasks. On numerical reasoning: 14.8%. On 171 of 750 tasks, none of the four models managed a pass. GPT-5.5: 25.7%, Gemini 3.1 Pro: 23.6%, Grok 4.3: 13.0%. LifeSciBench tests whether a model can do real scientific work, not whether it knows the terminology.
One task involves a gene therapy package for Duchenne muscular dystrophy. The model must assess whether the package holds up under FDA scrutiny: are the antibodies suitable for measuring the target protein, is the surrogate endpoint valid, does the external control distort the conclusion, is there enough data on safety and durability of effect. 25 criteria per task on average, 453 experts reviewed the tasks. This is not multiple choice.
Tacit Labs, OpenAI's partner on the project, starts with verification: in programming, AI has compilers and tests; in math, formal systems like Lean. In drug development, the final test is a trial in humans, and before that, years of indirect validation — cell experiments, animal studies, biomarkers, dose selection. The whole path is a chain of interdependent decisions: target → drug modality → molecule → assay → biomarker → patient group → clinical program. A mistake early on makes every later step expensive and useless.
Performance drops as soon as the task goes beyond text. GPT-Rosalind: 45.1% on text-only tasks, 28.1% on tasks with files or URLs. GPT-5.5: 29.9% versus 21.9%. On sequence and structure tasks, GPT-Rosalind scores 24.0%. 53% of the benchmark requires working with artifacts — charts, PDFs, tables, sequences, molecular structures, or links — and results are notably weaker there.
The benchmark has a limitation: the model gets a single shot, no clarifying questions, no follow-up experiments. Real science iterates — hypothesis, data, revision, next step. At 36.1% for the best model and 25 criteria per task, models have started to reason like scientific assistants but often fall short of a scientifically defensible answer.
750 tasks, 173 scientists wrote them, 79% require multi-step reasoning — and the best model cleared a third.
LifeSciBench directly benchmarks AI on tasks that include gene therapy regulatory reasoning, and our Gene Therapy glossary page provides the mechanism-level foundation needed to understand why the Duchenne FDA-package task is hard. The concrete numbers (36.1% best model, 14.8% on numerical tasks, 171 of 750 unsolved by any model) make the AI-in-biology claim falsifiable in a way Eternal Search readers can evaluate against the underlying science.