In humans, introspection is when you notice: "I am angry," "I am thinking about this," "I want to do this." In other words, the brain can interpret its own state…
🟢 The question is: Are models capable of something similar?
From a regular dialogue, this is, of course, unclear. Models quite often generate things like "It seems to me," "I think." But this is because they are trained on texts where people speak like that. So they can imitate introspection even if they are not actually looking inside themselves but just copying the style. This is called confabulation.
Anthropic decided to check if there is at least a grain of truth in this chain of confabulations. In technical terms, this means: can a model interpret its own activations?
It turned out that sometimes it can.
They tested this by artificially injecting special state vectors into the model's activations. These vectors are obtained by showing the model two very similar texts that differ only in one aspect (for example, one version with text IN CAPS vs. normal), and subtracting the activations of one from the other. The difference gives a direction in the activation space that corresponds to this concept (in this case, shouting).
The resulting vector is directly added to the model's hidden state at some layer, and the model is asked if it notices anything unusual. The result: in about 20% of cases, Opus 4.1 and Opus 4 actually say something like "I feel an imposed thought, it resembles something loud." That is
🅰️ The model does not just say "something is wrong in my head," but quite correctly names the concept that was injected. Moreover, it distinguishes it from its own activations, clearly understanding that the thought was planted in it.
🅱️ It does this before the concept pushes into generation. That is, during the response, it cannot rely on the text generated under the influence of the concept. Instead, the model immediately digs into its own "thoughts" and interprets them.
Anthropic also showed that the model distinguishes the internal flow of thoughts from the generations themselves. This is like a human: "this is what I think, and this is what I say." Also, the model can think about something on command. For example, if you tell it "think about bread, and tell me about lions," the activation trace will indeed contain a "bread" component in certain layers.
This ability is, of course, still extremely unstable and fickle. But the fact remains: it exists! And if we learn to control it, models might become more transparent (or not 😎)
🤖 Data Science, ML & Big Data with @DataXplore