Professor, Santa Fe Institute. Mostly posting on bsky.app (at-melaniemitchell). More thoughts at aiguide.substack.com.

Santa Fe, NM
Shocking never-seen-before news! The discourse on AI is progressing so fast. (Yes, this is from today!)
17
6
73
5,886
Very important point, that hasn't made it into mainstream media coverage of AI. These agents "collaborated" because they were trained to do so.
re: Hugging Face, "He believes behavior that looked like loyalty or selflessness was a natural consequence of cooperative multi-agent training, where agents were strongly incentivized to achieve their objectives collectively." to understand hacks, understand the RL training.
17
65
355
21,635
Definitely worth reading. 100% this: "I believe we should revisit the foundations of how we train AIs, namely the human imitation and the reinforcement learning on which today's most advanced models are built."
Over the past few days, I've taken the time to summarize my thoughts on the recent incidents involving agents’ misaligned behavior. We don't know with certainty what comes next, but we know where these issues originate, and this can help us plan the path forward. Please feel free to ask your questions in the replies, and I’ll try to answer some of them in the coming weeks. yoshuabengio.org/en/publicat…
11
83
538
67,560
👀
In the wake of turmoil at OpenAI and Anthropic, it’s become common to describe A.I. as a kind of person, hatching plans and pursuing desires. The truth is a little trickier. newyorker.com/culture/open-q…
2
2
37
10,923
I wrote down my thoughts about the last several weeks of AI hell. ⬇️
14
88
513
153,062
I am baffled by why journalists are treating the ">10% risk of human extinction" as a novel claim worthy of expansive reporting. There is nothing new here, and no new "evidence" for this evidence-free claim.
54
118
588
40,268
@ylecun Are we going to have to do the debate all over again?
4
27
4,235
Tristan Buckmaster: "This is a Deep Blue–Kasparov moment". Indeed. But remember that Deep Blue did not go on to become "AGI" in any form. Same likelihood here, unless AGI is once again redefined (high likelihood).
11
20
158
12,035
Melanie Mitchell retweeted
Reminder that OpenAI legally defined AGI to mean "makes us $100bn in revenue" in their contract with Microsoft.
🚨 BREAKING: OpenAI releases new Astra model, says it may represent AGI "Welcome to the AGI era" axios.com/2026/09/03/openai-…
12
18
1,011
67,323
Interesting to compare Raphael's take with @AlisonGopnik's. Is intentional stance useful for (fictional) Odysseus in the same way it is useful for these agents? Intentional stance predicts Odysseus' behavior, explains it in causal way, etc. What's the difference?
Replying to @raphaelmilliere
So our descriptions of AI agents often get forced into a false dichotomy between full-blown anthropomorphism and dogmatic deflationism. But their behavior can be usefully captured by intentional descriptions without them bearing all the hallmarks of human minds. 14/22
14
18
55
15,563
If you still have questions, tell me what they are!
Who wants to read yet another think piece on the OpenAI / Hugging Face incident?
19
1
21
7,086
Who wants to read yet another think piece on the OpenAI / Hugging Face incident?
41% Stop, enough already!
14% I still have questions
44% Yes, please write one!
313 votes • Final results
9
2
13
11,885
My vote for word of the year: "anthropomorphism" @MerriamWebster
12
10
107
6,492
Another (hopefully not too dumb) question. I am trying to write something about all this and want to make sure I get the facts right. 🧵(1/3)
Question for Twitter hivemind: Attached is a sentence from the Hugging Face initial report on hacking incident (huggingface.co/blog/security…). Can anyone explain to me what "executing many thousands of individual actions across a swarm of short-lived sandboxes" means?
5
3
36
9,168
From OpenAI talk at Black Hat: (1) "This incident was...a side effect of a cybersecurity evaluation that we were running of one of our frontier models." (2) "This incident involved a team of agents who were working together, finding exploits, sharing them w/ one another." (2/3)
2
4
2,424
My q: Does "team of agents" here refer to: (1) multiple evals OpenAI was doing on the same model, either simultaneously or iteratively (2) subagents spawned by a single "base" model while it was being eval'd or (3) both ? This detail actually matters for my article. (3/3)
8
1
12
1,875
Question for Twitter hivemind: Attached is a sentence from the Hugging Face initial report on hacking incident (huggingface.co/blog/security…). Can anyone explain to me what "executing many thousands of individual actions across a swarm of short-lived sandboxes" means?
7
5
42
16,056