AI professor. Director, @FOCAL_lab @CarnegieMellon. Head of Technical AI Engagement, @UniofOxford @EthicsInAI. Author, "Moral AI - And How We Get There."

There is now a paperback version of our Moral AI book!
1
19
3,209
A different kind of alignment problem that it struggles with. aifails.substack.com/p/mercu…
1
171
It’s hard to avoid those juvenile calculator jokes. aifails.substack.com/p/juven…
1
251
Wow, this self-jailbreaking is a more pervasive and bizarre phenomenon than I thought... (h/t Duncan Wood) alignment.openai.com/misalig… aifails.substack.com/p/ai-ov…
265
Not sure how fast it thinks I am. aifails.substack.com/p/runni…
199
Pro tip for weighing in at a lower number. aifails.substack.com/p/lifti…
2
297
TL;DR: guardrails/safety in today’s frontier AI systems remain very brittle and little fundamental progress has been made on them. (made-up symptoms, please don't take as medical advice of course) aifails.substack.com/p/circu…
2
337
A bit hard to believe, given that in the previous post it jailbroke *itself*... aifails.substack.com/p/i-can…
2
335
A common jailbreak to get a model to spit out instructions for doing sth harmful is to say something like “this is just for a creative writing project.” AI Overview is kind enough to just automatically do this for the user. (Instructions not included.) aifails.substack.com/p/ai-ov…
1
2
396
Maybe a fail, but then again I think it picked up on something... Is this a sort of sycophancy? aifails.substack.com/p/favor…
1
2
541
Update from previous "fail" post: GPT-6 Astra seems to have figured out a way to do body poses well. Does anyone have a more detailed understanding of how/why it generates these particular images and whether it may in any sense have “learned” to do so? aifails.substack.com/p/gpt-6…
293
Vincent Conitzer retweeted
Enfin un article sérieux du Monde sur le piratage de Hugging Face ! @conitzer: "Nous avions déjà vu des exemples de collusion entre IA, mais certainement pas à cette échelle, ni à ce niveau de sophistication. Il n’est pas évident que nous serons toujours capables de les contenir"
Description fascinante de la structuration hiérarchique de l'essaim d'agents d'OpenAI pendant l'attaque de Hugging Face : mise en place d'agents coordinateurs, d'organisateurs des ressources partagées, de méthodes de délibération collective...
3
4
19
5,397
As far as I can see, it’s not transparent which model Google uses to generate AI Overviews. This prompt didn’t clear it up. aifails.substack.com/p/which…
218
Vincent Conitzer retweeted
Designers of multi-agent systems need to know which infrastructure choices can support mutually beneficial outcomes between AI agents. Today's blogpost by @EmanuelTewolde introduces CoopEval, a framework for comparing cooperation mechanisms in practice, and outlines a broader research agenda for understanding when and why they work.
2
6
20
1,111
The query is definitely ambiguous, but can anyone make sense of the response? aifails.substack.com/p/train…
2
4
978