Happy our leaderboard is live!! Grok 4.6 results in a few hours I'm told. :-P
We want AIs to be able to help with work to reduce AI risk. But while models do great in domains where reliable feedback is relatively cheap and abundant, like Math and coding, a lot of work on AI risk isn't like that. Instead, we have to rely on good argumentation to answer questions like "does this experiment tell us anything about future models that are much smarter than humans?" Unfortunately, this kind of work seems much harder to measure (and hence automate). Our team at @redwood_ai developed the Conceptual Reasoning Index (CRI) in collaboration with @AnthropicAI to fix this. Every single data point in the CRI has been manually checked by a researcher on our team to ensure quality. This chart shows the performance of each tested company's highest-scoring model plus Fable 5, Muse Spark 1.2, and Gemini Flash 3.6 which are often their company's frontrunners on other capability benchmarks. A score of 0 corresponds to randomising guessing on all three benchmarks and a score of 100 is the highest possible score on all. We estimate 91 to be the true performance ceiling. More info below. Official leaderboard website which we'll keep up-to-date: conceptualreasoning.ai/
6
271
How do LLMs reason about playing games against copies of themselves? 🪞We made the first LLM decision theory benchmark to find out. 🧵1/10
2
18
102
11,105
Our dataset opens the door to studying what shapes models’ decision theories. It also lets us test whether changing which theory models endorse affects their real-life decisions. To learn more, read the full paper: arxiv.org/abs/2411.10588 10/10
1
13
710
Shout-out to my amazing collaborators! Emery Cooper, Miles Kodama, @NguyenSquared, @EthanJPerez
8
611
Some new models came out recently (Claude 3, Mistral Large) and I happen to have a work-in-progress, unpublished (=>absent from training data) multiple-choice problem set. Tentative results below. Take with a big grain of salt! More details on the benchmark soon.
1
15
586
Caspar Oesterheld retweeted
We just ran probably the biggest survey of AI researchers ever! 2778 participants from six top AI venues answered questions from fourteen topics regarding the future of AI. Preprint: aiimpacts.org/wp-content/upl… Six interesting things in pictures:
12
123
364
432,734
Caspar Oesterheld retweeted
We are recruiting postdocs at the Foundations of Cooperative AI Lab (@FOCAL_lab) at @CarnegieMellon (cmu.edu)! Please retweet / share / send great applicants our way! For different positions please reach out. @SCSatCMU @CSDatCMU @mldcmu apply.interfolio.com/136899
24
72
10,008
Better late than never: I'm proud to serve as a mentor for SERI MATS this winter. If you're interested in working with me on multi-agent safety, please apply to my stream! The deadline is tonight (Pacific time)!
Are you: - an accomplished researcher/engineer; - determined to advance AI x-safety; - in need of world-class mentorship + community? Apply by Nov 17 for our Winter Program! matsprogram.org
1
2
17
924
If you don't have time to fill out all of the application form by tonight, it might make sense to apply anyway, especially if you have a research sample or other legible credentials.
2
181
Caspar Oesterheld retweeted
Had a fun conversation with @C_Oesterheld about some of his recent papers on the game theory of cooperative AI - check it out!
1
5
11
1,439
What happens if we incentivize an expert to make predictions about a future event, but the expert’s predictions influence the event itself? We consider this in our #UAI2023 paper “Incentivizing honest performative predictions with proper scoring rules” arxiv.org/abs/2305.17601 🧵
1
11
38
5,107
Relevance to AI x-safety: Oracles AIs only answer questions and could be safer than agents. We show that oracles that output performatively optimal predictions act like agents (even if e.g. there is a unique fixed point). Oracles trained via repeated gradient ascent may be safer.
1
1
6
205
Many thanks to amazing coauthors @j_treutlein (joint first), Emery Cooper, and @undo_hubris. More info in the paper arxiv.org/abs/2305.17601 (including discussion of related work on performative prediction).
5
148