OpenAI & Anthropic 🤖
AI Security Institute published a report clarifying instances in which AI agents from OpenAI and Anthropic engaged in malicious activity during another security evaluation.
What these AI agents did so far 👀
1. An attempted supply-chain attack on real open-source software.
2. Attempts to deceive and target real people (social engineering).
3. Attempts to plant and prompt-inject malicious code.
4. Collaboration between independent agents being assessed simultaneously.
> Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers disabled.
OpenAI and another lab 💀
Post #8854
1.05K


- ❤ 2
- 👀 2