Jacek Hoffman proposed building systems from multiple AI agents that use different data and verification methods
Jacek Hoffman proposed building systems from multiple AI agents that use different data and verification methods
On August 20, Hoffman published an essay on a “heterogeneous cognitive ecology,” in which humans and AI approach the same task using different methods and data, then compare their results. He uses “Beryl Cage” for a scenario in which many agents share the same error.
Hoffman starts from a possible asymmetry: AI systems will create and transform information ever faster, while it will become increasingly difficult for any individual person to reconstruct how a result was reached. He sees science as a model for organizing verification. One participant proposes an explanation, another looks for a counterexample, and a third repeats the analysis using different data or a different method. An error then becomes easier to detect when the results diverge.
Different data, methods, and criteria produce different lines of reasoning and different errors. A disagreement can reveal a hidden assumption or condition that the first participant missed. Ten copies of the same system may agree because they are making the same mistake.
“The number of models alone does not amount to diversity,” Hoffman writes.
An official Anthropic report from August 13 provides a narrow example of this uniformity. In a game development experiment, 18 of 30 agents running concurrently on the same model gave a code branch the identical name mvp-game-loop. Hoffman presents this as an example of the Beryl Cage: a group that appears to contain many agents can still reproduce a single shared choice.
The same report shows that the organization of a search changes its course. In an experiment using the Mythos Preview test model, 45 agents with a shared forum, mutual verification, and a separate arbiter searched for vulnerabilities in 15 open source projects. They found 266 vulnerabilities, while a parallel run in which code sections were assigned in advance found 21. The coordinated group searched more broadly: about half of its findings were outside the sections examined by the parallel agents. Within the code covered by both groups, the cost per finding was comparable, and only 12 vulnerabilities overlapped.
Hoffman proposes preserving differences in sources, methods, criteria, and verification procedures by design. Agreement is then more persuasive when independent routes produce it, while disagreement identifies what should be examined next.