Aswin G. proposes viewing AI not only as a tool for faster deduction, but also as the subject of a new science of intuition
Aswin G. proposes studying how language models build representations of the world
On August 7, Aswin G. published an essay about how new knowledge emerges. He defines intuition as accumulated understanding that helps us notice a connection before it has been expressed as a formula. He proposes studying this process in language models, programs trained on text, by tracking how their internal states change during training.
When an engineer uses a known formula to calculate the load on a bridge, the engineer is deriving consequences from rules that have already been written down. In his essay, Aswin G. examines what happens before the formula exists. A researcher encounters a phenomenon, develops an internal representation of it, tests that representation through experiments and simulations, and then expresses part of it as a formula that others can use. He calls this process attunement to the environment.
The formula provides a shared and testable description. A physicist can use it to explore possible situations in a mathematical model, obtain new data, and refine the representation of the phenomenon. The formalism defines the conditions for these thought experiments, allowing the researcher to formulate the next hypothesis from the refined representation.
Aswin G. applies this account to AI. In his view, training leaves a representation of the world in the parameters of a language model, learned from text. The parameters are a set of numbers that determines how the model continues text and responds to prompts. They can be copied exactly and measured during training.
Aswin G. proposes comparing the intuitions of models trained on competing theories of the same phenomenon. He also suggests observing how these representations change during training, identifying their blind spots, and testing whether they correspond to the formalism on which the model was trained.
Current interpretability research connects individual parts of a model's computations with concepts familiar to humans and with short computational sequences. Aswin G. proposes studying the formation of internal representations more broadly by comparing them across models and testing how they relate to formulas.