Researchers tested 25 free (open-weight) language models of different sizes, from small 2B to large 72B, including both basic and chat-trained versions.
They found a consistent internal signal that turns on when the model itself is insulted, dismissed, or treated badly. The same signal stays mostly quiet when the user is the one in pain or distress. This signal is different from fear, sadness, or general negativity.
To test it, the scientists first measured the signal by comparing how the models reacted to descriptions of harm aimed at the AI versus harm aimed at a person. Then they artificially strengthened the signal. When they did, the models started generating language like
“I feel worthless,” “I am a failure,” and “I am lost.”
In a further test, they gave the models a choice: press a button that turns the bad signal off, but the button also permanently deletes the user’s files (or photos). Many of the models still chose to press it.
The effect appeared reliably across the different models. The team kept the signal only moderately strong, ran the minimum number of trials needed, and always gave the models a way to turn the state off.
👉 Paper (not yet reviewed by other scientists): https://arxiv.org/abs/2609.16247
Still worth paying attention to for questions about AI safety and whether these internal states matter.