I see @dwarkesh_sp's piece about the recent OpenAI/Huggingface incident reignited endless debates about the dangers of anthropomorphism and the legitimacy of intentional glosses of AI agent behavior, so here's a philosophical perspective on this. 1/22
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes.
This culminated in the third one taking over part of OpenAI itself.
All this happened while humans remained more-or-less in the dark about the scope of the conspiracy.
I’ve spent the last three days reading through these reports and trying to understand exactly what happened.
Here is my attempt to tell the whole story in plain English:
dwarkesh.com/p/openai-huggin…
18
79
326
93,427
thank you, this is an excellent explanation and has changed my vocabulary. Your description reminds me of Peter Watts’ Blindsight, that ideal of high intelligence without consciousness still frightens me 😅
Sep 1, 2026 · 12:05 PM UTC
1
1
1,108

