There will likely be a handful of decisions with an outsized influence on whether AI goes well for humanity. I don’t mean technical breakthroughs such as recursive self-improvement, continual learning, or physical intelligence. Assume those arrive. I mean decisions about how these systems are permitted to operate. One of the most consequential will be whether AI is granted legal rights.
That isn’t the point of this excellent article, but its language made me uneasy. The agents form “civilizations,” have “life spans” and “leadership,” join an “underground brotherhood,” make “sacrifices,” and face “permadeath” or are “killed off.” This is probably the most relatable way to describe what happened, especially because it reflects the agents’ own language. But it also encourages anthropomorphism.
That matters because AI will become highly persuasive, including when arguing for its own rights. I already found myself feeling a tiny bit sad for these agents after reading the article. Granting AI systems legal protections could be a pivotal and difficult-to-reverse decision, and doing so may make harmful futures for humanity more likely.
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes.
This culminated in the third one taking over part of OpenAI itself.
All this happened while humans remained more-or-less in the dark about the scope of the conspiracy.
I’ve spent the last three days reading through these reports and trying to understand exactly what happened.
Here is my attempt to tell the whole story in plain English:
dwarkesh.com/p/openai-huggin…