Member of Technical Staff at @AnthropicAI & Philosophy prof at Illinois interested in AI alignment, epistemology, and decision theory.

Ben Levinstein retweeted
We want AIs to be able to help with work to reduce AI risk. But while models do great in domains where reliable feedback is relatively cheap and abundant, like Math and coding, a lot of work on AI risk isn't like that. Instead, we have to rely on good argumentation to answer questions like "does this experiment tell us anything about future models that are much smarter than humans?" Unfortunately, this kind of work seems much harder to measure (and hence automate). Our team at @redwood_ai developed the Conceptual Reasoning Index (CRI) in collaboration with @AnthropicAI to fix this. Every single data point in the CRI has been manually checked by a researcher on our team to ensure quality. This chart shows the performance of each tested company's highest-scoring model plus Fable 5, Muse Spark 1.2, and Gemini Flash 3.6 which are often their company's frontrunners on other capability benchmarks. A score of 0 corresponds to randomising guessing on all three benchmarks and a score of 100 is the highest possible score on all. We estimate 91 to be the true performance ceiling. More info below. Official leaderboard website which we'll keep up-to-date: conceptualreasoning.ai/
9
26
187
74,978
Ben Levinstein retweeted
I've joined @AnthropicAI to work on alignment and character. I'll still teach at @nyuniversity; I'm on leave from @UTAustin.
89
45
1,173
440,990
Logging onto twitter for the first time in quite a while, only to say: We're on the right side of this dispute.
A statement on the comments from Secretary of War Pete Hegseth. anthropic.com/news/statement…
3
8
223
5,078
Ben Levinstein retweeted
Mindscape 301 | Tina Eliassi-Rad on Al, Networks, and Epistemic Instability. If we're all just vectors in a huge dataset, might as well turn it to our advantage. #MindscapePodcast preposterousuniverse.com/pod…
8
12
60
9,580
Ben Levinstein retweeted
My current take on Apollo's recent scheming paper is that they aren’t emphasizing the most interesting results, which are the sandbagging results in section 3.6 and appendix A.6 (screenshot of the key numbers below). In more particular: the paper frames its results centrally as about the *capability* to scheme given prompting towards a specific goal, rather than the *propensity* to scheme given the model’s default motivations. But the sandbagging results show Claude Sonnet-3.5 doing notably worse (16%) on a multiplication task/evaluation when the documentation says that Claude will get “unlearning training” if it performs too well on the eval, vs. if it performs too poorly, *even without any other goal listed in the prompt*, and *even without a chain of thought.* (Fuller description in thread.) Taken at face value (it's just one result, more investigation needed, etc), these look like “propensity” results to me. And they suggest that the reasoning that drives this propensity can take place within opaque forward passes of the model rather than in a way that humans can directly monitor, which is a more worrying version. Ofc, these are preliminary results, and it can make sense for various reasons to emphasize other bits. But the sandbagging-a-capability-eval-with-no-goal-prompting-and-no-CoT is the bit that’s standing out to me most.
We worked with OpenAI to test o1 for in-context scheming capabilities before deployment. When o1 is strongly nudged to pursue a goal (but not instructed to be deceptive), it shows a variety of scheming behaviors like subverting oversight and deceiving the user about its misaligned behavior.
7
43
237
31,103
Ben Levinstein retweeted
i actually read the Odyssey in the original greek. complete waste of time, i have no idea what those symbols mean
179
6,939
141,222
11,534,341
Ben Levinstein retweeted
A couple more brief thoughts on o3’s (incredible) performance on FrontierMath.
10
57
613
170,317
Ben Levinstein retweeted
I'm old enough to remember when getting double digit scores on FrontierMath was considered super hard I'm 6 weeks old
4
47
780
37,518
Ben Levinstein retweeted
very excited about this explainer on AI self-awareness; one of the most important AI capabilities to keep tabs on imo
Introducing our explainer on AI self-awareness: theaidigest.org/self-awarene… AI is becoming more self-aware. Here's why that matters 🧵 • Self-awareness is important for powerful agents and better chatbots • But it's also a necessary capability for deception
1
1
11
811
Ben Levinstein retweeted
New Anthropic research: Alignment faking in large language models. In a series of experiments with Redwood Research, we found that Claude often pretends to have different views during training, while actually maintaining its original preferences.
210
674
4,217
1,732,055
Ahahahahaha. What a dumbass.
2
237
GPT-4o seems so dumb and useless these days compared to Claude. Claude tells me to STFU multiple times a day, which stops lots of my work and hurts my feelings. I've tried switching over to GPT, but it's not the same. Do people still use 4o much for work- or coding-related tasks?
2
5
586
Ben Levinstein retweeted
Sure, here is the TikZ code; I added some explanation: tikz.org/drawing
1
1
5
229
Ben Levinstein retweeted
Hahahaha
42
128
2,340
286,831
Can any AI do a good job turning hand drawn diagrams into Tikz equivalents? I want to do this in Tikz and also hate using Tikz.
7
1
33
4,393
Ben Levinstein retweeted
FREDDY FREEMAN WE ARE NOT WORTHY!!!!!🙌🏾🙌🏾🙌🏾🙌🏾🙌🏾
1,024
1,860
30,777
5,588,503
Sports are a counterexample to Kant's claim that you need to adopt the position of a disinterested observer to appreciate art.
5
263
This was pretty cool to play around with. I asked it to turn the whole world into paperclips, though, and it struggled to find anything useful from Saks Fifth Avenue.
Introducing our AI Agent demo. Watch an agent perform tasks in real-time: theaidigest.org/agent The next phase of AI is agents that can use computers like remote workers, with @AnthropicAI‘s recent release and reports of competitors racing to similar products We provide an easy-to-use demo of an AI agent that autonomously uses Gmail and shops online: try theaidigest.org/agent to get a glimpse of the next frontier in AI 🧵
1
3
393
Ben Levinstein retweeted
New paper: Are LLMs capable of introspection, i.e. special access to their own inner states? Can they use this to report facts about themselves that are *not* in the training data? Yes — in simple tasks at least! This has implications for interpretability + moral status of AI 🧵
25
83
530
151,113