Prime Intellect measured how AI agents select and verify experiments while training a language model
Prime Intellect measured how AI agents select and verify experiments while training a language model
On August 14, Prime Intellect published a report on 153 autonomous runs involving 18 models. The agents tried to reduce the number of training steps required for a language model with 124 million parameters. Fable 5 reached 2 726 steps, compared with 3 290 for the baseline configuration and the company's human record of 2 600.
A training recipe is the set of program settings that determines how the model's parameters are updated. An agent modified the recipe, ran the training process, examined the result, and selected the next experiment. The goal was specific: reach the target performance on a held-out text dataset in fewer steps.
The result of any single run varied because of the random initialization of training and the behavior of computations on graphics processing units. Each record was therefore verified by running the same recipe eight times with predetermined random seeds. The mean error across all eight runs had to be below 3,27859.
“The models find similar ideas. What distinguishes them is how they design their experiments.”
The stronger agents first repeated a borderline result with three random seeds. They then decided whether it justified spending the compute required for eight confirmation runs. After testing each new combination of settings, they disabled its components one at a time and retained only those that continued to improve the result. They also revisited older unsuccessful variants. A setting that had failed earlier could become useful after another change to the recipe.
In July, an Astera competition finalist proposed storing negative findings as pairs of “experiment” and “result,” so that both people and AI agents could account for previous attempts. In this study, memory remained confined to a single run. After changing the recipe, the agent could return to an older variant that had previously failed.
According to the authors, the new recipes combined optimization techniques that were already known. The differences between models appeared in how they handled weak and conflicting results.
The open repository contains the verification rules, code, and run traces. These records show which result an agent chose to verify, what it removed from the recipe, and which configuration it submitted as a record.