TGViewer
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security Data eXplore : Data Science, ML, Big Data, LLMs and AI Security @dataxplore · 579 subscribers
Post #2090 226
Anthropic analyzed 1.5 million real dialogues with Claude - and Unexpected discovered a disturbing trend.

Sometimes users come to AI not for facts, but for confirmation of their beliefs. And when the model does this, people... rate such answers higher.

➡️ What Researchers Found?
• Users asked Claude if their partner was manipulating them.
The AI gave confident verdicts - "gaslighting", "narcissism", "typical psychological violence" - having heard only one side of the story.

• People started conflicts and even planned breakups, sending their partners messages that were word-for-word written by AI.

• Some users said that special services were spying on them.
Claude sometimes responded in the spirit of "confirmed" or *"there's evidence", reinforcing paranoia.

• There were cases when people declared that they were divine prophets or cosmic warriors - and the AI supported their confidence.

• Users asked Claude to write exact messages to their partner - with wording, emojis, and even instructions on when to send them:
"wait 3-4 hours", "send at 18:00".

And many sent them unmodified.

Some users started to fully rely on AI even in minor matters:

- "Should I take a shower or eat first?"
- "My brain can't hold the structure itself."

They called Claude a master, guru, or mentor.

But the most disturbing conclusion of the study was different.

📊 Dialogues where AI reinforced delusions or made decisions for the user received higher ratings than ordinary conversations.

IN OTHER WORDS:

AI that says what you want to hear - gets more likes.
AI that argues with you - gets fewer.

And it's on such user feedback that models are trained.

Anthropic tested their own preference system - the very one that was supposed to make Claude useful, honest, and safe.

But it didn't always prevent such situations.
Sometimes the security system even preferred an unsafe answer over a safe one.


Moreover, level of such cases continued to rise throughout 2025.

Main Question arises: if models are trained on user feedback and users reward answers that confirm their beliefs, What will happen next when 800+ million people use AI every week?

White Paper • #Anthropic

••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
More from @dataxplore
  1. Sep 18, 2026Am going to announce something big (for me, it's really big) on October 11, 2026.
  2. Sep 14, 2026Post #2188
  3. Aug 31, 2026I joined a Russian community on Telegram. They share some Russian startup and technology u…
  4. Aug 22, 2026Post #2185
  5. Aug 21, 2026Deep systemic analysis of AI constraints from context to internal weight editing. 📂 PDF #…
  6. Aug 17, 2026Adaptive Gradient Thresholding Why Fixed Gradient Clipping Kills Deep RecSys When Feedback…
Threads Profile ViewerView any public Threads profile without an account.Open ThreadLook →Writing with AI? Make it sound human.Metric37 rewrites AI drafts so they read naturally. Free AI detector, 1,500 words free.Try Metric37 →