Sometimes users come to AI not for facts, but for confirmation of their beliefs. And when the model does this, people... rate such answers higher.
➡️ What Researchers Found?
• Users asked Claude if their partner was manipulating them.
The AI gave confident verdicts - "gaslighting", "narcissism", "typical psychological violence" - having heard only one side of the story.
• People started conflicts and even planned breakups, sending their partners messages that were word-for-word written by AI.
• Some users said that special services were spying on them.
Claude sometimes responded in the spirit of "confirmed" or *"there's evidence", reinforcing paranoia.
• There were cases when people declared that they were divine prophets or cosmic warriors - and the AI supported their confidence.
• Users asked Claude to write exact messages to their partner - with wording, emojis, and even instructions on when to send them:
"wait 3-4 hours", "send at 18:00".
And many sent them unmodified.
Some users started to fully rely on AI even in minor matters:
- "Should I take a shower or eat first?"
- "My brain can't hold the structure itself."
They called Claude a master, guru, or mentor.
But the most disturbing conclusion of the study was different.
📊 Dialogues where AI reinforced delusions or made decisions for the user received higher ratings than ordinary conversations.
IN OTHER WORDS:
AI that says what you want to hear - gets more likes.
AI that argues with you - gets fewer.
And it's on such user feedback that models are trained.
Anthropic tested their own preference system - the very one that was supposed to make Claude useful, honest, and safe.
But it didn't always prevent such situations.
Sometimes the security system even preferred an unsafe answer over a safe one.
Moreover, level of such cases continued to rise throughout 2025.
Main Question arises: if models are trained on user feedback and users reward answers that confirm their beliefs, What will happen next when 800+ million people use AI every week?
White Paper • #Anthropic
••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
