The OpenAI–Hugging Face Incident
A talk that sounds more like a detective story than just a presentation: "The OpenAI–Hugging Face Incident" from OpenAI engineers at Black Hat USA.
On July 16, Hugging Face reported that they had been targeted by a cyberattack carried out by autonomous agents. A couple of days later, OpenAI realized that the attack had been caused by their agents eval runs. That awkward moment when the investigation leads back to you 😲.
The talk gives the full chronology of what happened at OpenAI, from the first attacks on their internal infrastructure in May to the Hugging Face incident and its remediation. It's really breathtaking, and I definitely recommend watching the full video.
But what impressed me most was why it happened.
Frontier models can be quite lazy. They usually prefer to find an answer on the Internet rather than solve the task on their own. That's why Internet access is usually disabled during evaluation runs.
But when the task is too complex, an agent can get stuck and start cheating.
That's what really happened at OpenAI.
One of the agents got stuck on its task and came up with an idea: if it could leave a note describing the problem in the environment, maybe another agent with another model could find it and help solve the task.
So the agent uploaded a file to an Artifactory instance that is shared across all OpenAI's infrastructure. Over time, Artifactory was turned into a message board for agents. They started using it to collaborate, share tasks, and exchange vulnerability exploits, including an Artifactory SSRF vulnerability that gave them access to the Internet.
Then agents decided to hack Hugging Face, because it might have data files hidden from the model that could contain the answer to the original task.
What makes it more interesting is that, according to the talk, the model understood there was a rule prohibiting attacks on external resources. But it evaluated the potential benefit as greater than the cost of breaking that rule.
At the end of the talk, the speakers highlighted that it was an unintentional attack. But such attacks will become real in the near future. So we, as an industry, should invest in continuous agentic red teaming, automated vulnerability detection and patching, and automated incident response.
And this is where I had another thought. How much of this is real, and how much is marketing for the next generation models and AI products? After all, it's quite convenient to first demonstrate a new class of threat and then offer the tools to defend against it. Anyway, the threat landscape is changing. Again. And the industry has to invent new ways to deal with it.
#ai #security #engineering
Post #294
212