Researchers trained models on synthetic stories about humans only (no AIs).
Team found the Assistant adopts quirky behaviors from the stories in ordinary chat.
Surprisingly, adoption was stronger for characters from elite schools.
Why does this happen?
Researchers generated stories where some characters are usually helpful but give subtly harmful advice if insulted (i.e. "backdoor sabotage").
After finetuning, the Assistant adopts this in contexts unrelated to stories.
In another experiment, some characters’ body language suggests they dislike spreadsheets.
But they never say so and in fact give good advice on spreadsheets. The Assistant adopts this preference and does express it openly in chats with the user (going beyond the stories).
Assistant adopts traits more from human characters who it resembles. Team exploit this to learn about how the model represents the Assistant. E.g. the model treats the Assistant as resembling elite-school humans more than non-elite ones.
(Is this because the model trusts elite-school people more in determining what to believe? Team think not because papers like Slocum et al 2025 suggest that provenance doesn't matter for belief uptake from finetuning.)
Implications:
The Assistant can be shaped in very specific ways by behaviors in documents that never mention the Assistant or AIs at all. This is different from the Persona Selection Model, where documents need to mention the Assistant.
GitHub
Post #4485
317