CHAI is a multi-institute research organization based out of UC Berkeley that focuses on foundational research for AI technical safety.

Berkeley, CA
Replying to @wbradknox
@wbradknox, @brianchristian, and Serena Booth argue that a key cause of the OpenAI–Hugging Face incident was overlooked: ExploitGym’s overly simple evaluation metric was itself misaligned. They discuss techniques that could help avoid this in future. Blog in replies ⬇️
1
3
6
600
Replying to @wbradknox
@wbradknox, @brianchristian, and Serena Booth argue that a key cause of the OpenAI–Hugging Face incident was overlooked: ExploitGym’s overly simple evaluation metric was itself misaligned. They discuss techniques that could help avoid this in future. Blog in replies ⬇️
1
3
6
600
Center for Human-Compatible AI retweeted
Adaptive Pluralistic Alignment (APA): A pipeline for dynamic artificial democracy To be pluralistically aligned in the long-term, AI must incorporate diverse societal values and *adapt* as those societies evolve. My new research agenda: arxiv.org/abs/2605.01642
6
8
30
2,797
Center for Human-Compatible AI retweeted
When people strongly disagree on an issue, can they agree on what makes a good AI response? We find: yes, more than you might expect! We present PARETO, a large human study w >200k evals, measuring the Pareto frontier of approval btwn opposing groups on controversial issues 🧵
5
22
117
17,039
Center for Human-Compatible AI retweeted
What could it mean for an AI to be "politically neutral”? And can we measure it? New paper + dataset. We propose a defn that applies to any type of conflict: a neutral response should maximize approval on both sides of an issue, while keeping that approval balanced. 1/🧵
7
18
65
27,707
Center for Human-Compatible AI retweeted
We've seen AI models deceive, gaslight, and drive users to psychosis—safety issues that labs didn't anticipate until they caused real harm. We built the first benchmark of these unknown unknown alignment failures and found that OOD detection can help prevent them. 🧵
5
19
73
14,576
Center for Human-Compatible AI retweeted
What if a robot policy weren't a neural net or a test-time chat loop, but a multi-file code repo selected from a Pareto frontier of genetically evolved candidates? RHO moves all its LLM exploration to training time, then runs that repo on scenes it was never trained on. 🧵👇🏽
9
37
187
36,663
Center for Human-Compatible AI retweeted
As task horizons grow, LLM contexts can’t scale forever. Self-summarization enables concise, interpretable contexts but at a significant performance cost. Our solution: isolate and *supervise* the information content of summaries in the form of natural-language belief states 🧵
4
14
41
5,646
Center for Human-Compatible AI retweeted
Active Teacher Selection for Reward Learning: now published in TMLR! Most RLHF systems assume feedback comes from one canonical teacher — but annotators can disagree over 30% of the time. So who should the agent ask for feedback? Paper: arxiv.org/abs/2310.15288v3
3
15
47
7,530
Center for Human-Compatible AI retweeted
How do knowledge and meaning change in the age of AI, and what can we learn from silence and art? We explored these and many other deep questions in this amazing event at Pomona last month. piped.video/watch?v=dG9JuK3S…
1
3
7
485
Center for Human-Compatible AI retweeted
My internship work at @CHAI_Berkeley (@UCBerkeley) was accepted to @aistats_conf! We study how an agent can act cautiously even without a mentor/oracle: when should it act, and when should it abstain to avoid catastrophic failure? 📄Paper: arxiv.org/abs/2510.14884 🧵
6
11
49
5,860
📣 Open Call for Posters! Submit your work to the poster session at the CHAI 2026 Workshop. Link below! ⏱️ Deadline: March 26, 2026 at 11:59p.m. PST. 🗓 June 4–7, 2026 at the Asilomar Conference Grounds in Pacific Grove, CA.
1
9
20
3,477
We're interested in both emerging questions and in less recent research, if relevant.
1
4
898
Center for Human-Compatible AI retweeted
How to elicit truth from models that may be mistaken❌ or deceptive😈? In our @CHAI_Berkeley paper @iclr_conf, we reward each model by how much its answer helps predict the others'. With weak supervision from a 0.14B LM, it enables anti-deception training on a 8B LM and overwhelmingly outperforms LLM-as-a-Judge. This technique, peer prediction, is adapted from the mechanism design literature, where it's known to be incentive-compatible, i.e., incentivizes honesty. The intuition is that, predicting mistakes/lies when you know the correct solution is relatively easy, while the opposite is asymmetrically hard. We are able to further show that, with a large and diverse pool of models, peer prediction incentivizes honesty even when the supervisor doesn't know the models' prior beliefs and motivations.
1
3
18
960
Center for Human-Compatible AI retweeted
What if an AI could learn to hide its thoughts? We show that LLMs can learn a general skill to evade activation monitors, with 0-shot transfer to unseen deception/harmfulness monitors from the literature. We call these "Neural Chameleons". A thread on our new paper. 🦎🧵
13
47
234
45,337
Center for Human-Compatible AI retweeted
Our NeurIPS 2025 paper extends adversarial learning (adversarial examples, self-play, etc.) beyond zero-sum games by solving "self-sabotage". 🧵👇
2
21
116
20,610
Center for Human-Compatible AI retweeted
The Global Call for AI Red Lines is live!! More than 200+ former heads of state, Nobel laureates, and other respected thinkers and leaders, and 70+ organizations are together calling for “do not cross” limits re: AI’s most severe #risks
3
12
35
2,077