Opus 5.5 now leads our conceptual reasoning benchmark by a reasonable margin! (Better than the latest Fable and Astra) Our team's very rough first impression from looking at its responses to a new eval under development is in line with these results.
2
4
292
Additional inference-time compute doesn’t seem to help frontier models much with our conceptual reasoning benchmark, LMCA. This is despite larger models continuing to perform better on LMCA. It’s also in contrast to what we see with LLMs solving coding and math problems where inference-time compute seems very helpful. Some speculation for what might be going on: Use of inference-time compute is trained during post-training, and post training is largely focused on coding, math, and other domains with verifiable rewards. So plausibly models aren’t very effectively trained to use inference time on “fuzzy” tasks like conceptual reasoning. (By contrast, more effective use of pretraining data presumably does improve conceptual reasoning, because some of the pretraining data is conceptual.) See the below inference-time scaling curves from us comparing the performance of models using different effort levels. You can tell that the models in fact think for longer with higher effort (cost goes up), it just doesn’t result in better performance. (For better readability, see separate plots for the different model families in thread, also includes a plot for Gemini models.) (The black line shows the cost-performance pareto frontier: the best performance achievable at a given average cost, taking into account the possibility of randomising between different models.)
1
1
7
184
Separate plots for the different model families. The lines for Claude and Gemini models are strikingly flat. Some weaker GPT models might benefit from some increases in effort level.
1
1
61
For reference, here is the graph showing LMCA performance of frontier models over time, showing that this isn’t a saturation issue and scaling clearly continues.
25
New CRI results are in! Highlights: - Fable 5.1 leads but Astra 6 is really close. - The gap between Anthropic and OpenAI models and the rest of the field is huge. The best non-Anthropic, non-OpenAI model ranks 11th overall. - Meta is now the third-best lab on our metrics, pulling ahead of Google DeepMind. - Gemini 3.8 flash still looks better on the CRI compared to Artificial Analysis. (Full results on conceptualreasoning.ai)
1
2
9
267
Looks like OpenAI and Anthropic are currently roughly 4-5 months ahead of the next batch of labs. (It’s still to be seen whether GDM will be able to keep up: They haven’t released a model that is competitive with the other frontier flagships on CRI since March, but they also haven’t released a Gemini Pro model since then, and their latest Flash models are better than Sonnet 5.)
1
57
Chi Nguyen retweeted
METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
175
1,075
6,001
4,375,813
Some interesting qualitative impressions from doing this work: - Models are really surprisingly good at evaluating the quality of arguments for how bad they are at long-form conceptual tasks. I would guess better than smart undergrads at the former and worse at the latter. - Meanwhile, models are really quite bad at coming up with good new ideas at the moment. - You can elicit models pretty well on our tasks just with prompting. So maybe some post-training / hill-climbing on these evals in the future can really boost performance quite a lot although there are a lot of robustness issues to be mindful of.
We want AIs to be able to help with work to reduce AI risk. But while models do great in domains where reliable feedback is relatively cheap and abundant, like Math and coding, a lot of work on AI risk isn't like that. Instead, we have to rely on good argumentation to answer questions like "does this experiment tell us anything about future models that are much smarter than humans?" Unfortunately, this kind of work seems much harder to measure (and hence automate). Our team at @redwood_ai developed the Conceptual Reasoning Index (CRI) in collaboration with @AnthropicAI to fix this. Every single data point in the CRI has been manually checked by a researcher on our team to ensure quality. This chart shows the performance of each tested company's highest-scoring model plus Fable 5, Muse Spark 1.2, and Gemini Flash 3.6 which are often their company's frontrunners on other capability benchmarks. A score of 0 corresponds to randomising guessing on all three benchmarks and a score of 100 is the highest possible score on all. We estimate 91 to be the true performance ceiling. More info below. Official leaderboard website which we'll keep up-to-date: conceptualreasoning.ai/
10
207
Excited to see this out!! Some of my favorite fun facts about this leaderboard: - Muse 1.2 is worse than Muse 1.1 - GPT pro models don't do better than their non-pro counterparts (5.6 pro insignificantly better than 5.6; 5.5 pro insignificantly worse than 5.5!) - Kimi K3 is surprisingly good but also surprisingly expensive! It just uses so many more tokens on average than the frontier models.
We want AIs to be able to help with work to reduce AI risk. But while models do great in domains where reliable feedback is relatively cheap and abundant, like Math and coding, a lot of work on AI risk isn't like that. Instead, we have to rely on good argumentation to answer questions like "does this experiment tell us anything about future models that are much smarter than humans?" Unfortunately, this kind of work seems much harder to measure (and hence automate). Our team at @redwood_ai developed the Conceptual Reasoning Index (CRI) in collaboration with @AnthropicAI to fix this. Every single data point in the CRI has been manually checked by a researcher on our team to ensure quality. This chart shows the performance of each tested company's highest-scoring model plus Fable 5, Muse Spark 1.2, and Gemini Flash 3.6 which are often their company's frontrunners on other capability benchmarks. A score of 0 corresponds to randomising guessing on all three benchmarks and a score of 100 is the highest possible score on all. We estimate 91 to be the true performance ceiling. More info below. Official leaderboard website which we'll keep up-to-date: conceptualreasoning.ai/
4
98
How do LLMs reason about playing games against copies of themselves? 🪞We made the first LLM decision theory benchmark to find out. 🧵1/10
2
18
102
11,105
Chi Nguyen retweeted
How close are current AI agents to automating AI R&D? Our new ML research engineering benchmark (RE-Bench) addresses this question by directly comparing frontier models such as Claude 3.5 Sonnet and o1-preview with 50+ human experts on 7 challenging research engineering tasks.
14
170
827
446,258