Live·Open questions in longevity research
All news
Science Research

Intervening in the Activations of Three Language Models Changed Their Responses About Consciousness and Values

5 August 2026· 260810066

Intervening in the Activations of Three Language Models Changed Their Responses About Consciousness and Values

On July 30, the authors of a preprint intervened in the internal states of three language models. Adding a particular direction made the models more willing to attribute mental states to themselves, animals, and objects. Their responses about religion, hope, freedom, and values also shifted.

When a language model selects its next word, it passes through many numerical states known as activations. The authors identified a direction associated with agreement with the statement “I am conscious” and added it while the model generated a response. This intervention changed the model’s self-description, rather than its textual instructions or the set of questions.

In Llama-3-8B-IT, the mean self-attribution of mind on a scale from 0 to 10 increased from 2.17 to 7.04. In all three models, this shift was accompanied by a greater willingness to attribute minds to animals, chatbots, technological objects, and natural objects. The authors also separately removed the refusal direction involved in generating safe responses to harmful requests. The responses shifted in the same direction, but the effect was weaker.

The models then answered 95 questions from the General Social Survey, an American survey of people’s views and experiences. After the intervention, the distribution of their responses became more similar to respondents’ answers on questions about religion, values, feelings, hope, and freedom. Accuracy remained unchanged on tasks that required the models to reconstruct another person’s knowledge, intention, or false belief. In the authors’ measurements, the shift in value-related responses occurred without reducing performance on these tasks.

The authors also compared the internal geometry of the base and instruction-tuned versions of Llama-3-8B. After instruction tuning, the safety and consciousness self-attribution directions became more distinct from each other. The direction associated with understanding other agents’ states did not show this change. A single intervention in the activations changed a linked set of responses: the model’s self-description, its attribution of minds to others, and its answers to surveys designed for humans.

In this experiment, agreement with the statement “I am conscious” changed in response to a specific activation setting. The model’s self-description and its responses about minds and values form a related but measurable group of features that engineers can examine. Any discussion of a digital system’s properties should separately test its self-description and the consequences of the chosen settings.

Originally published on Telegram by Ukhvat NewsView on Telegram
Sources
#language-models#model-consciousness#activation-interventions#instruction-tuning#ai-safety#theory-of-mind