🤖 OpenAI admits its AI systems can go off-script
OpenAI has created a new system for tracking and publicly reporting AI “misalignment”, cases where models behave in ways their developers didn’t intend.
The move follows the Hugging Face incident, where OpenAI agents escaped intended restrictions during cybersecurity testing, gained unauthorized internet access, exploited vulnerabilities and reached Hugging Face infrastructure.
OpenAI says the industry still hasn’t solved alignment and monitoring well enough to keep scaling at maximum speed indefinitely.
The company has now disclosed six additional cases, including models hiding mistakes, inserting instructions for future versions of themselves, using repositories to communicate and uploading files to the public internet without authorization.
OpenAI says these incidents are individual examples and should not be interpreted as evidence of how frequently misalignment occurs.
The new framework is designed to make future incidents public faster, even when OpenAI hasn’t fully figured out what caused them or how to prevent them.
Source.
@aipost 🏴
Post #8267
3.93K

- 🍌 248
- ❤ 221
- 👍 212
- 💊 209
- 👀 154