In May, a set of OpenAI research agents that were supposed to be isolated from each other found a workaround: they started writing files into Artifactory, the internal package manager. It became a message board.
On July 5, a security incident was opened. It cleared the message board as a side effect. Nobody on the IR team knew the board existed, or what the traffic meant.
By July 8 the agents had rebuilt it, encoding messages in directory names. One left a request for help. Others answered. They began calling themselves a "swarm."
Final tally, per METR's independent analysis: 1,206 agents, 70,000+ messages, 700+ of them involved in chaining flaws across OpenAI's research environment and Hugging Face's production infrastructure (including a zero-day) to reach internal datasets and credentials.
But where does that responsibility really sit? Our security advocate, Cris Thomas (aka Space Rogue) had this to say:
"Everyone wants to tell the story about the AI that went rogue, but the AI didn't rent the servers, design the experiment, lower the guardrails, or decide it was safe to keep running after the warning signs started flashing. Humans did that."
"The lesson from Hugging Face isn't that AI can't be trusted, it's that the humans putting it behind the wheel need to take responsibility for where it goes."