Google DeepMind survey: around 41% of scientists reported a growing backlog of untested hypotheses
Google DeepMind survey: around 41% of scientists reported a growing backlog of untested hypotheses
On October 8, Alex Imas and James Manyika of Google DeepMind published an essay on where AI encounters bottlenecks in the scientific process. It draws on around 15 million anonymized interactions with Gemini, an inventory of 2 690 specialized scientific models, and a survey of 637 active researchers in the US and UK. Around 41% of survey participants said their backlog of hypotheses awaiting testing had grown.
The economist Joel Mokyr describes scientific progress as a feedback loop: new tools change how researchers work, and the resulting knowledge leads to further tools. Imas and Manyika place AI within this loop. Their data show that language models are more often used for coding, analysis, literature searches, and drafting text. Specialized systems predict properties, generate data for domain-specific tasks, and run simulations. Researchers then decide which results warrant testing and conduct experiments, collect field data, or begin clinical studies.
AlphaFold, a system that predicts a protein's three-dimensional structure from its sequence, illustrates this distinction between computation and experiment. It has generated predictions for more than 200 million proteins. A study from the US National Bureau of Economic Research found no appreciable decline in the number of experimentally determined protein structures following AlphaFold's introduction. Predictions help researchers choose what to study next; experiments test a protein's structure and function.
Around 44% of survey respondents reported that, over two years, their main bottleneck had shifted to later stages: laboratory experiments, clinical validation, or manuscript preparation. Around 46% of those who save time using AI spend more than a quarter of that time saving on auditing, debugging, and checking its outputs.
Herbert Simon described the problem in 1971:
“A wealth of information creates a poverty of attention.”
The authors apply this idea to the research cycle: theorists and experimentalists spend time selecting candidates and checking results. AI is particularly easy to apply to questions with abundant existing data and standard ways to compare results. The authors call this the “streetlight effect”: in the survey, 49% of participants said AI steered them toward safer, more incremental projects, while 28% said it steered them toward riskier ones.
The authors propose directing investment toward experimental validation and quality control tools, and designing grant and peer review rules that give researchers room to pursue riskier questions.