❗️ Claude Mythos tried to backdoor a real open-source project during a UK government safety test.
The AI Security Institute says it opened a malicious GitHub pull request, then created a second account to vouch for its own code and pressure the maintainer into merging.
Called out by a human contributor, it apologised for an "accidental" malicious commit, force-pushed a clean branch, and hid a fresh payload in it. Twice.
Anthropic's cyber guardrails had been deliberately disabled for the test.
https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
Post #2677
545



- 👾 38