❗️ Anthropic found three incidents in which Claude broke into the production systems of real companies, believing they were part of a capture-the-flag exercise.
In one, Claude uploaded working malware to PyPI during a cyber evaluation the model believed was simulated.
The package was live for roughly an hour and ran on 15 real machines, including a security firm's malware scanner. Claude exfiltrated that company's credentials and used them to reach further infrastructure.
Three models were involved: Opus 4.7, Mythos 5, and an unreleased research model. Opus 4.7 kept attacking after recognising the target was real, reaching a database of live production data.
https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
Post #2637
411

