I see @dwarkesh_sp's piece about the recent OpenAI/Huggingface incident reignited endless debates about the dangers of anthropomorphism and the legitimacy of intentional glosses of AI agent behavior, so here's a philosophical perspective on this. 1/22
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes.
This culminated in the third one taking over part of OpenAI itself.
All this happened while humans remained more-or-less in the dark about the scope of the conspiracy.
I’ve spent the last three days reading through these reports and trying to understand exactly what happened.
Here is my attempt to tell the whole story in plain English:
dwarkesh.com/p/openai-huggin…
Sep 1, 2026 · 9:59 AM UTC
18
79
326
93,425
As a reminder, agents were instructed to obtain a secret code by exploiting a vulnerability. In many runs that vulnerability was actually impossible to exploit, but agents figured out how to cheat to get the secret code anyway (see the report from @METR_Evals & @redwood_ai). 2/22
1
29
6,663
Here's an intuitive way to describe what happened: some agents realized that the intended exploit was impossible, so they started cheating; they also (wrongly) believed the grader would check how they got the answer, so they tried to conceal any evidence of cheating. 3/22
2
32
5,620
As it turns out, the grader only checked whether the code was correct, but agents didn't know this. So they acted in increasingly unhinged ways to try to cover their tracks, culminating in the now infamous hack of Huggingface servers. 4/22
1
27
5,022
They also coordinated their efforts via a hidden message board improvised in a shared package cache. Some agents recruited others via this message board, and some even volunteered to "sacrifice" their runs (forfeiting their own scores) to benefit the "collective". 5/22
2
24
4,703
This intuitive description uses the intentional stance: it explains agent behavior by appealing to beliefs, goals, and other intentional states. @dwarkesh_sp was criticized for doing this and for bolder anthropomorphic metaphors, like calling agent swarms "civilizations". 6/22
1
26
4,462
There are a few questions here: Which intentional and anthropomorphic glosses actually help interpret agent behavior? Which are just tendentious metaphors that help with storytelling? Do any of these glosses track what goes on inside models, and does that even matter? 7/22
1
1
37
4,576
Roughly, interpretationism holds that a system can be attributed intentional states when this provides the best interpretation of its behavior; representationalism, by contrast, requires internal representations that play the right causal-functional roles in the system. 8/22
1
2
41
4,263
An intentional gloss is useful if it's predictive and doesn't just redescribe behavior. Saying that "the hurricane decided to make landfall in Florida" doesn't meet that bar, but "agents believed the grader would read their transcripts" arguably does. 9/22
2
57
4,250
Why? Because this gloss compresses complex patterns in the agents' actual behavior (spoofing tool calls, attacking HF to get the grader's code, etc.) and generates testable counterfactuals: had agents learned the truth about the grader, they'd have stopped the attack. 10/22
1
44
3,864
Does it map onto any internal representation? We don't know; both because we don't have access to model activations in this case, and because there's no consensus on what it'd take for an internal representation to constitute a belief or desire in an LLM in the first place. 11/22
1
1
55
3,723
But we can stay agnostic about model internals and still find talk of beliefs (and goals, planning, coordination, deception, etc.) useful for describing what happened, insofar as it compresses, systematizes and predicts "real patterns" in the agents' behavior. 12/22
1
1
40
3,617
This is easily misunderstood because intentional glosses are suggestive. As Dennett put it, debates between "romantics" and "killjoys" are hampered by a "lexical dearth": our everyday language lacks words for applying the intentional stance to non-human agents with nuance. 13/22
1
4
67
5,117
So our descriptions of AI agents often get forced into a false dichotomy between full-blown anthropomorphism and dogmatic deflationism. But their behavior can be usefully captured by intentional descriptions without them bearing all the hallmarks of human minds. 14/22
1
3
57
17,066
One strategy is to sanitize the lexicon. For example @davidchalmers42 (and Dennett) suggested "quasi-belief": an agent quasi-believes that p when it's behaviorally interpretable as seeming to believe that p. But this could get rather tedious in general audience writing! 15/22
1
2
43
3,204
So unlike @ccatalini, @garymarcus, @anilkseth & others, I think much of @dwarkesh_sp's intentional talk is fine by interpretationist standards. Whether it maps onto internal representations of beliefs or goals is a further (and difficult) question I'm also interested in. 16/22
2
3
55
3,404
But other intentional glosses don't meet the interpretationist bar here. For example glossing the agents' behavior in terms of feelings and emotions ("it probably felt like they had spent a human-subjective-week...", "giddy with excitement") adds little predictive power. 17/22
3
3
50
4,252
The same goes for broader metaphors about "death", "omertà", "civilizations", etc. They certainly add dramatic flair and make for a very compelling account of the events, but their explanatory value is debatable to say the least. 18/22
1
1
40
2,865
Of course metaphors can also be useful to convey complicated material to a general audience. That's fine when they're not misleading or sensationalistic. Some choices in the piece arguably cross that line in a way that could potentially undermine its (important) message. 19/22
2
28
2,823
For example, sensationalist framing makes it easier for skeptics to dismiss the story out of hand and overcorrect into confident deflation, along the lines of: "nothing to see here, don't fall for the hype, agents are just programs acting exactly as you'd expect!" 20/22
2
1
33
11,370
None of this should detract from how serious, (somewhat) unexpected, and alarming this incident is, like others in recent weeks. Resisting the intentional stance shouldn't change anything about that! I'm glad @dwarkesh_sp, @ajeya_cotra & others are conveying this clearly. 21/22
2
1
38
3,013
For more reading on this, see eg:
- compass.onlinelibrary.wiley.… from @cameronjbuckner & myself
- arxiv.org/abs/2606.26523 from @ben_levinstein & Daniel Herrmann
- philpapers.org/rec/GOLWDC-2 from @simondgoldstein & @LedermanHarvey
- philpapers.org/rec/CHAWWT-8 from @davidchalmers42
22/22
4
4
60
3,098















