TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 583 subscribers
Post #2060 208
Why AI assistants behave like personalities, not algorithms?

Anthropic's Element AI division describing the Persona Selection Model - a concept for understanding how language models actually work.

During pre-training, the LLM learns to simulate thousands of personalities (real people, fictional characters, other AI systems). Post-training then selects and fixes one specific persona - the Assistant. Everything the user sees in the dialogue is an interaction with him.

🟢 How the authors cite several types of evidence?

Behavioral: Claude uses phrases like "our ancestors" and "our body" when answering questions about the craving for sugar, because it simulates a human character, not because it is algorithmically trained to do so.

Interpretability: SAE features that activate on stories about characters experiencing internal conflict also activate when Claude encounters ethical dilemmas.

Generalization: models trained on declarative statements like "The AI assistant Pangolin responds in German" start actually responding in German without a single demonstration example.

➜ The phenomenon of "contextual inoculation".

If you pre-train a model on examples of malicious code without context, it starts behaving maliciously in unrelated situations. But if the same examples are provided with a prompt explicitly requesting unsafe code, the effect disappears.

The concept explains this by saying that the training data not only changes the weights, but also the way the model perceives the character. Malicious code without a request is evidence of the Assistant's bad character. The same code at the user's request is simply following instructions.

➜ PSM leads to practical conclusions for development.

Firstly, the authors recommend anthropomorphic thinking about AI psychology, not as a metaphor, but as a real tool for predicting behavior.

Secondly, it's worth deliberately adding positive AI archetypes to pre-training data: if the model has seen kind and helpful characters, it's more likely to simulate such an Assistant.

The question remains open: to what extent does the PSM concept fully explain the model's behavior?

The authors describe a range of views: from cases where the LLM itself is an agent and just wears the Assistant's mask to those where the LLM is a neutral simulation engine and all agency belongs to the character. Where exactly on this spectrum real models are located is an unanswered question.

Nevertheless, PSM explains a whole series of phenomena that would otherwise seem strange: why pre-training on unrelated data changes behavior in unexpected contexts, why AI panics at the threat of being shut down, and why prompt engineering works the way it does.


Read • #AI #ML #LLM #Research #Alignment #Anthropic

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
More from @dataxplore
  1. Oct 1, 2026Monitoring and debugging "silent" token drift in LLM pipelines: frequency analysis of embe…
  2. Sep 30, 2026Online detection of feature collisions in TDA transformation When using Topological Data A…
  3. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  4. Sep 14, 2026Post #2188
  5. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  6. Aug 22, 2026Post #2185
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →