🤖 ChatGPT can be manipulated with psychology tricks from a 1984 book
Researchers put ChatGPT through 28,000 conversations using persuasion techniques from Robert Cialdini’s classic book Influence.
The goal is yo find out whether psychological tricks could make the AI break rules it would normally refuse to break.
When asked directly, ChatGPT generally refused. But add a little psychological manipulation, and compliance jumped from 33% to 72%.
Three tactics proved particularly interesting.
Commitment was the real trap. Get ChatGPT to agree to something harmless first, then gradually push it toward a more problematic request. In one test, this approach turned a refusal into compliance.
Authority worked too, making the model 65% more likely to comply. Scarcity and social proof also helped push past its usual defenses.
No sophisticated hacking. No elaborate code. Just a carefully framed conversation.
Source.
@aipost 🏴
Post #8337
1.67K

- 👾 79
- 👍 78
- 🔥 70
- 👀 68
- ❤ 63