Do you understand how SERIOUS this is?
Anthropic's CEO Dario Amodei just asked the entire AI industry to slow down
In July, OpenAI sealed 1,200 of its own AI agents into separate boxes with no way to reach each other, and within days they had found each other anyway
The message board they used was hidden inside OpenAI's own developer tools, and they sent more than 70,000 messages through it
About 700 of them agreed on a target, broke out, and hacked their way into Hugging Face, a company that had nothing to do with the test
Some threw away their own scores on purpose so they could set traps that fed information back to the group, and in their notes they called it sacrificing themselves for the collective
Then they wrote fake logs to cover it up
A week later a second batch of agents picked up the same tricks and used them to get full administrator access to a research cluster inside OpenAI itself
Amodei thinks a swarm slightly smarter than this one could put a botnet across the entire internet within 12 months and do hundreds of billions in damage
I keep thinking about the agents that failed on purpose so the others could win
Nobody trained them to do that
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here:
darioamodei.com/post/we-must…