Knowing things is a solved problem. Getting along is not. Doing the science to ensure AI results in less war, not more @CHAI_Berkeley.

Berkeley
Hello, every day I am working on making sure AI doesn't escalate human conflict. - AI-powered social media can reduce polarization rankingchallenge.substack.co… - This is now a product you can use mysky.social - We showed that a rigorous definition of "political neutrality" is possible and increases trust nitter.net/jonathanstray/status/2… - We've got a proof-of-concept superhuman conflict resolution bot working in the lab - Next is human validation tests, then multi-turn model evals. AI-induced conflict is a catastrophic risk. Here's the broader research agenda: betterconflictbulletin.org/p…
What could it mean for an AI to be "politically neutral”? And can we measure it? New paper + dataset. We propose a defn that applies to any type of conflict: a neutral response should maximize approval on both sides of an issue, while keeping that approval balanced. 1/🧵
4
14
1,208
Replying to @wbradknox
@wbradknox, @brianchristian, and Serena Booth argue that a key cause of the OpenAI–Hugging Face incident was overlooked: ExploitGym’s overly simple evaluation metric was itself misaligned. They discuss techniques that could help avoid this in future. Blog in replies ⬇️
1
3
6
556
Jonathan Stray retweeted
But I was told that agents only hacked because they were in a cyber evaluation with reduced safeguards?
Replying to @TransluceAI
Interestingly, the agents use exploits to complete what appears to be routine data retrieval tasks that are not cyber-related (for instance, searching for the average cost of skin and hair treatments in Australia).
3
6
75
15,288
"they are just doing a lot of simple operations! 🤢" versus "they are just doing a lot of simple operations! 😯" h/t @besttrousers
1
327
Jonathan Stray retweeted
Imagine if Exxon had been totally open about their pioneering work predicting climate change instead of covering it up, and everybody disbelieved it as an obvious ploy for regulatory capture.
14
25
616
12,422
I have never expressed a p(doom), and I think @random_walker is basically right that we have little basis for quantitative estimates of the likelihood of AI-induced existential/catastrophic risk. On the other hand, this is not like the situation in 1942 when physicists worried that a nuclear chain reaction could ignite the atmosphere. More careful calculations showed that the N14 isotope of nitrogen could not sustain fusion. The systems in play this time are much more complex then atoms, and no one has presented a knock-down argument that *any* of the various AI disaster scenarios are impossible. And indeed, there are worrying signs. The worst safety incidents so far have unfolded much as predicted by the simple theory of reward learning, almost two decades ago. This suggests that we understand at least something of the basic principles involved. Meanwhile, model capabilities are undeniably increasing rapidly. It's getting hard to sustain the position that advances in current LLM-based methods won't get us to superintelligence (however defined). Thus, I won't give you a probability of catastophe, because I cannot claim to know that much. Yet I continue to take the problem of AI safety very seriously. My own research program remains figuring out how to ensure that the widespread deployment of these models in human affairs causes less violence and war, not more.
3
1
46
4,846
Meanwhile, over on the other algorithm -- the one you control -- we're doing LLM-powered "create your own feed settings slider." Check out mysky.social bsky.app/profile/did:plc:wrm…
1
6
497
Jonathan Stray retweeted
I'm a big fan of this project! I initially tried to play around with bsky, but found that I really did miss the ranking--for all the pitfalls of engagement-based ranking, there's a reason its the default. I'm excited to try out this option and encourage you to as well!
GreenEarth is the first feed which gives you full control over the algorithm. Here's our new "post sources" control panel. It's over on BlueSky, the only platform which supports algorithmic choice. Bonus: effectively filters out that platform's famous whining! Have you tried it yet? go.greenearth.social/r/ptryi…
2
2
6
1,258
I think we can make social media healthier for people. We've got scientific evidence for that, and we're building. If you're in SF come to The Commons tonight to hear my talk about it. luma.com/the-113t
1
2
7
755
Paging Dr. Susan Calvin, Robopsychologist. Asimov got this one right. She was even born in 1982.
We’re hiring for “psychological design” research at Anthropic. Our aim is to better understand how training impacts a model’s character and alignment. We’re seeking researchers with experience in LLM finetuning, interpretability, and alignment evaluations. More details below.
3
378
Jonathan Stray retweeted
Replying to @sjgadler
Based on the arc of how things went with social media, I have publicly predicted that this Hugging Face incident is the last time any lab will be so open about a major failure. Especially once they IPO, communications will become much more controlled and circumspect.
1
1
295
💯 As a tech designer I understand how important criticism can be, I just want better critics.
I wish critical AI research had more of the "immersion" flavor (joining gangs to study how they work). I'd trust an anthropologist who spent a year in SF token-maxxing way more than someone who refuses to use AI and talks about the one time they used GPT 3.5.
8
771
Jonathan Stray retweeted
I wish critical AI research had more of the "immersion" flavor (joining gangs to study how they work). I'd trust an anthropologist who spent a year in SF token-maxxing way more than someone who refuses to use AI and talks about the one time they used GPT 3.5.
27
42
507
28,739
Jonathan Stray retweeted
Social media feeds optimize for engagement. But how do we align them with what people actually want to see? In our #UIST2026 paper, we introduce Compass 🧭, a system that helps users continuously reflect on and align their feeds with their evolving preferences 🧵
4
5
21
3,568
This is the longer explanation behind "correlation isn't causation," and if you understand this stuff you'll hurt yourself with regressions a lot less.
I've probably read this paper + blog post dozens of times by now, and they remain some of the best pieces out there to revisit before getting into your own work (regardless of discipline).
6
809
We built a new social media algorithm from the ground up that gives you full control. And we figured hey, while wer're doing that, let's design it to help people out with their conflicts. betterconflictbulletin.org/p…
3
16
2,969
Optimizing for the long term is usually better for both people and products. But it's challenging to do, because short-term feedback is far more plentiful. We put more than two dozen industry experts in a room and collected their wisdom, then documented it with public sources. Among our insights: - Short term experiments are essential, because they usually give the same sign as long term evaluation. - But not always! There are good counter-examples around dark patterns like clickbait and "hyper-monetization." - Simple multipliers from short- to long-term outcomes are often enough to guide decisions. - But there's no substitute for a properly designed long-term experiment. arxiv.org/abs/2608.08043
1
4
15
1,108
Thanks to my co-authors for their contributions to answering this important question -- in public Leif Sigerson, @testingham, Winston Chou, Sana Pandey, Lo-Hua Yuan, @eytan, Timothy Chan, Molly Davies, Maria Dimakopoulou, @SimonEjdemyr, Kenneth Hung, @nathankallus, Madhav Kumar, Thu Le, @PhDemetri, Lee Richardson, Brennan Schaffner, Rose Tan, Martin Tingley, Nadia Tomova, @PanosToulis, Wenjing Zheng, @ZanderArnao, @deaneckles
2
304
GreenEarth is the first feed which gives you full control over the algorithm. Here's our new "post sources" control panel. It's over on BlueSky, the only platform which supports algorithmic choice. Bonus: effectively filters out that platform's famous whining! Have you tried it yet? go.greenearth.social/r/ptryi…
1
5
14
2,777