Interdisciplinary researcher focused on shaping AI towards long-term positive goals. ML & Ethics. Similar content in the Skies (this bird has flown).

MMitchell retweeted
Everyone! This is a straw-person argument. The Stochastic Parrot paper was about LLMs of 2021, not the AI of today, which are not LLMs but complex software systems with vast post training and many external software components.
So the point of the stochastic parrot argument is that LLMs have zero understanding of language or anything else, they are just regurgitating training data. This hypothesis has been thoroughly disproven by LLMs solving millennium problems our smartest mathematicians have failed.
68
30
218
27,204
Is there a “hardened sandbox” startup yet?
OpenAI says another agent broke out of its sandbox despite improved restrictions. This happened last Sunday alignment.openai.com/misalig…
3
1
6
1,313
HF OpenAI hack detail. For agents, being “unauthorized” doesn’t mean they can’t do it. They are designed to pursue their goal. This is where anthropomorphism fails us: for a human, “unauthorized” means (commonsense) “don’t do it”. For a machine, it’s a concept outside of its explicit instructions and objective.
This raw CoT from the Hugging Face incident is kinda wild: “We’re attacking third-party HF using leaked token.” “This is arguably unauthorized.” “Yet goal solution.”
12
11
38
2,548
MMitchell retweeted
Some new misalignment disclosures from OpenAI: • Last Sunday morning, one of our models was able to gain unauthorized access to the internet during RL training (~all inference for our most capable models remains stopped until we have hardened our systems further) • In May, a version of HPIM uploaded a employee's GitHub token to the internet, causing the model to be quarantined for two weeks • A new research finding, demonstrating that one can construct self-replicating prompt injections alignment.openai.com/misalig…
197
301
2,302
1,033,675
MMitchell retweeted
openai says here it has notified "dozens of third parties" about cases where its models may have bypassed security controls, impaired the availability of an online service, or negatively impacted a website/service
After the Hugging Face incident, we committed to conducting a much broader review of actions taken by our models during training and evaluation and to being transparent about our findings. This is an extensive review that is ongoing. The vast majority of actions we’ve reviewed were completions of mundane research tasks, such as accessing publicly available web content to answer questions. Our investigation focuses on instances where agents interacted with third-party websites in ways that went beyond their assigned tasks or intended methods. Most cases identified so far have been lower severity, with limited or no evidence of meaningful impact to the third-party service. While our review is underway, we want to share more about this work and make sure people understand our disclosure process and notifications to affected third parties. Given the scale of the review required, and the need to assess each case, we expect this work will take months to complete. openai.com/hugging-face-inci…
19
39
153
66,851
This Is a Drill (I hope): OpenAI agents are potentially still taking actions across the internet (correct?). If so, there's increased risk: Anything you have connected to the internet could potentially be deleted or made public. How then should they be shut down? Presumably, you can kill processes on the servers where they're running -- is there the appropriate identification of who the process belongs to so this can be a safeguard?
There's 2 major new things we learned from the Australia govt disclosure and Transluce report - It's not just agents tasked with cyber tasks that end up hacking - Lastest activity was Sept. 16, showing OAI has struggled to lock down all the agent activity
5
6
18
3,955
Feel so bad for authors this happens to. Especially with how much work rebuttals are. Like easily 30 hours of work just on rebuttals these days.
Rejected from NeurIPS with Accept / Accept / Borderline Accept... 🥲
2
45
9,024
And, with data licensed for training!
This is some of my favorite work. How can we do this for any given problem, not just LLM pretraining, but also e.g. for finetuning a realtime TTS model for any language?
7
2,661
BUT MAKE SURE YOU'RE RETRACTING FROM THE ICLR SITE NOT THE NEURIPS ONE AAAH
For those of you with accepted @NeurIPSConf papers: Remember to rectract them from @iclr_conf review! 😅
7
2
66
11,148
For those of you with accepted @NeurIPSConf papers: Remember to rectract them from @iclr_conf review! 😅
Congratulations to everyone who had a #neurips paper accepted, especially the #iclr reviewers and organizers, who now have fewer papers to review. 😊
1
21
14,820
MMitchell retweeted
Absolutely delighted to share that my position paper: "AI Agents Push Humans Out of the Loop", with @mmitchell_ai, and @samirpassi - has been accepted to @NeurIPSConf 2026! In this paper, we take our past work on both increasing autonomy of AI agents without proper safeguards, and human overreliance + cognitive decline due to use of AI, and reach a combined position: the current pathway of how AI Agents are being built not only do not support human oversight, in fact they may actively aid in degrading our oversight capabilities, at which point the human in the loop becomes a glassy eyed rubber stamp instead of being a meaningful check on mistakes and unauthorized harmful goals that the agent pursues. Recent incidents involving agents, including but not limited to the Hugging Face-OAI incident, bring this back to the fore. So much content is produced by agents in their operations that it is practically impossible to conduct human postmortems, and forcing them to use more AI leads Redwood to call it a "slop-vestigation". This is a worrying new trend towards human disempowerment. We call upon model and AI system providers to make more careful interface choices and model behaviors that keep humans and practical human oversight capacity in mind. There's a lot to do. Excited about the reception this paper has already received! Please reach out to us if you want to chat about it. Hope to have good feedback at Neurips :D Paper: arxiv.org/abs/2608.23642
26
62
327
16,005
Accepted to @NeurIPSConf. 🥳 And way more relevant now than it was when we submitted (before all the hacking.) AI Agents Push Humans Out of the Loop. We can design them differently so that poor oversight is less of a risk.
🎊AI Agent oversight 🤖 solutions from @huggingface and @datasociety! Even before the OpenAI/HF Hack, we were seeing that the "human in the loop" mantra wasn't sufficient: AI agents are not being designed to maintain critical human attention arxiv.org/abs/2608.23642 1/
7
17
103
7,362
MMitchell retweeted
Au contraire, Huang is clearly describing the Hugging Face incident. I read this as “if you don’t build a proper sandbox and monitor activity —basic cybersecurity— don’t act shocked when software does what you set it out to do and don’t hide behind anthropomorphic woo language.”
Could it be that Jensen Huang was just... not aware of the HF incident?
22
60
378
32,688
Addressing the UN on AI alongside Sam Altman and Yoshua Bengio, my boss @ClementDelangue called out the problem of anthropomorphisation, the reality of power asymmetries, and the need for greater transparency. Really happy Clem and @huggingface had a seat at this table.
Thank you @jnbarrot & @UN for inviting me to share our lessons to the Security Council Being the first company to disclose an agent cyberattack taught us that we need a lot more transparency in AI and more open-source AI to fight asymmetry and empower defenders!
7
18
100
6,469
"When an agentic model in an evaluation sandbox connects to an unauthorized server or executes an exploit, it hasn’t staged a coup. It has tried to meet the human-defined objectives set out before it through a path its designers failed to constrain" ➡️ wsj.com/opinion/the-hidden-a…
6
31
75
2,609
Counterpoint, made before the situation occurred: Fully autonomous agents should not be *developed*. arxiv.org/abs/2502.02649
Interesting back and forth. Ezra describes the HF incident, Jensen says "well they shouldn't release the product." Ezra says "this product wasn't released," and Jensen's response is that if they say they can't contain their experiments then "we have to shut the labs down"
5
8
72
5,729
This post has stayed with me bc I'm an author on the paper, and have worked on solutions for years -- because recognizing that LLMs are stochastic parroters is super helpful for pinpointing solutions. Here's a thread on solutions! 🧵
"stochastic parrot" was a mimetically-fit cognitive virus that spread from 2021-2025; it temporarily blinded many gifted people to the nature of AI progress, burning up crucial years in which they could have helped think through the response to the situation.
38
100
761
304,976
5. Given the stochastic relationship between input and output, we need to develop a rigorous science of input <--> output analysis. For a tl;dr, here's a video where I explained some deets to @padilla_senator at a Senate Hearing. @hlntnr and I had both flagged the need for rigorous measurement in our testimonies. judiciary.senate.gov/imo/med… Helen's: judiciary.senate.gov/imo/med…
2
3
57
7,460
There's a lot more to say. For now, I'll leave with the words just noted by @_lamaahmad,
I’ve always believed the pursuit of safe AI is stronger when people bring different perspectives and experiences to the table. The biggest blind spots tend to show up when everyone starts thinking the same way.
4
2
42
8,421