LLMs accelerate the research cycle but do not produce verifiable results
Two coding errors forced Christopher Meiklejohn to withdraw results from a study conducted with an AI agent
On August 15, Christopher Meiklejohn published an essay titled “Science at LLM Speed”. He created SetScope, a program that identifies songs in concert recordings. AI agents, programs that use large language models to complete tasks with multiple steps, quickly produced the code and analysis. Two errors changed what the results meant.
Meiklejohn asked an agent to study improvisation in concert recordings. The agent was to collect the files, prepare the program's training data, and produce reports. In the first analysis, some performances appeared in both the training data and the validation set. The duplication was obscured by different copies, excerpts, and encodings. The performance metric was unusually strong because the program encountered familiar material during validation. After an audit, Meiklejohn removed this analysis.
In the second attempt, the methodology explicitly ruled out identifying improvisation from fixed 90-second segments. The program's calculations still used the old value: 90 seconds. Based on those reports, Meiklejohn published two posts and invited people he knew to participate in a listening study. He later withdrew the posts, deleted the second notebook, and stopped the study.
“The form did not create the errors. It made them harder to notice.”
Both errors changed what the analysis actually measured. In the first case, the estimate of the program's performance depended on the composition of the validation set. In the second, improvisation was defined using a rule that contradicted the methodology. The AI agent quickly produced code, plots, and explanations for one person to review. Meiklejohn concluded that data provenance and the measurement rule need to be checked separately.
The MDA system selects the next experiment where competing explanations make different predictions. Such an experiment can rule out a specific conclusion.
In the spring, the ICLR machine learning conference applied a similar procedure to submitted manuscripts. An automated system identified suspicious references, which people then checked against bibliographic databases and web search results. At least three people reviewed every confirmed case, and manuscripts containing fabricated references were withdrawn from consideration.
“What unfavorable answer could it return? What precise claim would that answer rule out?”
In research conducted with agents, each successive step should rest on data, a measurement, or an experiment capable of changing the conclusion.