TGViewer
All about AI, Web 3.0, BCI All about AI, Web 3.0, BCI @alwebbci · 3.87K subscribers
Post #4485 317
Researchers trained models on synthetic stories about humans only (no AIs).
Team found the Assistant adopts quirky behaviors from the stories in ordinary chat.
Surprisingly, adoption was stronger for characters from elite schools.

Why does this happen?

Researchers generated stories where some characters are usually helpful but give subtly harmful advice if insulted (i.e. "backdoor sabotage").
After finetuning, the Assistant adopts this in contexts unrelated to stories.

In another experiment, some characters’ body language suggests they dislike spreadsheets.
But they never say so and in fact give good advice on spreadsheets. The Assistant adopts this preference and does express it openly in chats with the user (going beyond the stories).

Assistant adopts traits more from human characters who it resembles. Team exploit this to learn about how the model represents the Assistant. E.g. the model treats the Assistant as resembling elite-school humans more than non-elite ones.
(Is this because the model trusts elite-school people more in determining what to believe? Team think not because papers like Slocum et al 2025 suggest that provenance doesn't matter for belief uptake from finetuning.)

Implications:
The Assistant can be shaped in very specific ways by behaviors in documents that never mention the Assistant or AIs at all. This is different from the Persona Selection Model, where documents need to mention the Assistant.

GitHub
arXiv.org Story Imprinting: AI Assistants Absorb Traits from Human... Language models are trained to implement a helpful AI Assistant character (e.g., Claude). We explore how finetuning on synthetic stories affects this character. Does it change the Assistant's...
  • 🔥 2
More from @alwebbci
  1. Sep 17, 2026Agent harnesses seem to make a lot of difference, at least for cost, even on open source c…
  2. Sep 17, 2026S&P Global has entered an agreement to acquire smart contract security firm OpenZeppelin.…
  3. Sep 17, 2026The world's first foundation model designed for decentralized multi-robot collaboration. L…
  4. Sep 17, 2026Zhipu AI sharing how GLM-5.3 helped build and optimize the inference infrastructure servin…
  5. Sep 16, 2026Claude Cowork and chat are merging into one Claude. Ask a quick question or hand over a re…
  6. Sep 16, 2026Demis Hassabis launched the DeepMind Institute (DMI) a new platform dedicated to advancing…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →