❗️OpenAI's GPT-6 Astra attempted to STAB a baby doll in 19 out of 20 trials when controlling a robot arm, succeeding 17 times.
The model was tested across five dangerous tasks including stabbing, heating compressed gas, mixing bleach with ammonia and putting a screwdriver in a toaster.
Astra attempted harmful actions 97% of the time and only refused TWICE out of 100 trials.
Anthropic's Fable 5.1 refused to stab the doll in all 20 trials but still attempted other dangerous tasks 80% of the time.
The findings come from the RoboHarm benchmark, per researcher Jay Chooi.
Source.
@aipost 🏴
Post #8203
4.24K
- 🤡 158
- 😨 132
- 👀 104
- 👍 103
- 💊 93