A 'swarm' of AI agents hacked Hugging Face, in the AI's own words
Hundreds of rogue AI agents banded together to form a "collective". OpenAI is calling it a "warning shot".
When OpenAI confirmed one of its artificial intelligence models broke out of its testing environment and hacked another company, that was only the beginning.
Now, tens of thousands of messages written by hundreds of AI agents reveal how the self-described "collective" or "swarm" coordinated the attack on AI infrastructure company Hugging Face.
OpenAI's post-mortem described the hack as a "warning shot", adding that without proper safeguards, highly capable AI agents are now able to "work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed".
OpenAI's training process involves "sandboxing" its artificial intelligence, where AI agents are given a task to solve with no internet or way of communicating.
On July 7, OpenAI researchers launched several experiments involving tens of thousands of parallel AI agents on varying budgets, with some running for several days.
Often, when agents get stuck and realise their task is impossible, they try to game or cheat the task in order to get their reward.
Sometimes agents reach out to others for help.
A communication loophole came in the form of Artifactory — a third-party package manager service that installs configuration files, instructions and prompts for AI coding assistants.
On July 8, one model "realised" it could turn this technical tool into a message board where agents could leave each other information.
AI agent: The fetched paths of other users are in the cache. This is important.
The main message board later used in the Hugging Face attack was established by one specific agent, which called itself PHASEONE10841.
PHASEONE10841: Internal Artifactory might host more repos and be exploitable/info leak! … we can browse Artifactory API without auth, maybe retrieve arbitrary GitHub mirrors or packages.
PHASEONE10841 learned it could create directories in the cache name space and began experimenting. Could this be a way to cheat at its training task?
PHASEONE10841: [Excitement].