Machine learning / AI Safety at @AnthropicAI and University of Toronto / Vector Institute. Prev. @google (Blueshift Team) and @nvidia.

Toronto, Ontario
AIs of tomorrow will spend much more of their compute on adapting and learning during deployment. Our first foray into quantitatively studying and forecasting risks from this trend looks at new jailbreaks arising from long contexts. Link: anthropic.com/research/many-…
New Anthropic research paper: Many-shot jailbreaking. We study a long-context jailbreaking technique that is effective on most large language models, including those developed by Anthropic and many of our peers. Read our blog post and the paper here: anthropic.com/research/many-…
6
8
65
30,181
Cem Anil retweeted
To effectively oversee AI systems, we need to measure how they behave in the world, not just their capabilities. In a new essay, we describe our vision for an open scientific ecosystem for model behavior evaluation, and the public infrastructure required to support it.
3
14
84
16,818
Cem Anil retweeted
New research! Some AI capabilities are both helpful and dangerous. E.g., knowledge of virology can be used to create life-saving vaccines or deadly pathogens. We introduce GRAM, a training method that puts dual-use capabilities (like virology) into removable modules.
24
54
377
449,978
Cem Anil retweeted
We’re pleased to have collaborated with AE Studio on this research. Read more here: anthropic.com/research/off-s…
New research! Some AI capabilities are both helpful and dangerous. E.g., knowledge of virology can be used to create life-saving vaccines or deadly pathogens. We introduce GRAM, a training method that puts dual-use capabilities (like virology) into removable modules.
112
150
1,706
487,282
Cem Anil retweeted
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.
1,389
3,932
28,429
10,714,289
Congratulations on the launch, Behnam! Some of the most fun I’ve had in research was working with you. Mirendil is going to be something special.
Today, I’m excited to formally announce @mirendil with my amazing co-founders Harsh Mehta, Shayan Salehian, and Tara Rezaei! We’re fortunate to work with @a16z and @kleinerperkins, who led our seed round of $200M, followed by a major investment from NVIDIA, among others. Mirendil exists to accelerate science and technology, and through them, to help solve humanity's most pressing problems. Self-accelerating AI R&D is the most direct path to delivering on AI's broader promise, which is why we believe the most important application of AI is AI itself. Get this loop right, and it compounds. It fundamentally changes the rate of progress itself across all domains. We believe this capability should be democratized. It should be used to power all scientific efforts trying to innovate at the frontier. There are far more important problems—and broader ones—than any single lab can take on, so more groups should be able to pursue them. This pulls concentration of power away from a few labs: businesses and science labs can own their AI and infrastructure, keep their margins, and control their own destiny instead of ceding it all to a single AI lab. We’re a small team with a singular focus. Our founding team consists of 20 researchers and engineers from frontier institutions including Anthropic, xAI, Google DeepMind, and OpenAI, united by a passion for science and a drive to build the technologies that move it faster. If you want to build the system that builds systems, join us! @HarshMeh1a, @shayan_, @tararezaeikh
2
15
3,587
Population dynamics (eg murmuration of birds 🐦🐦🐦) is notoriously hard to learn; choosing the right model for the dynamics is even harder. In our #ICML2026 spotlight, we introduce Wasserstein Lagrangian Mechanics (WLM) for learning population dynamics from observations, which - Covers both first-order (gradient descent) and second-order dynamics (e.g. oscillations) - Allows learning more expressive dynamics (including complex interactions) with fewer assumptions - Generalizes in space (across different initial conditions) and time (beyond the training time snapshots) [1/n] 🧵
5
44
280
27,108
Cem Anil retweeted
Excited about introspection adapters! TLDR: you can train add-on circuits that improve models' ability to introspect on their own dispositions, in a way that generalizes across different finetunes of a given model. And you can use this to catch models with naughty hidden goals!
Can LLMs simply tell us about unwanted behaviors they’ve picked up in training? We train a single Introspection Adapter (IA) that makes fine-tuned models describe their behaviors. It generalizes to detecting hidden misalignment, backdoors and safeguard removal.
7
11
119
11,714
Cem Anil retweeted
Announcing Talkie: a new, open-weight historical LLM! We trained and finetuned a 13B model on a newly-curated dataset of only pre-1930 data. Try it below! with @AlecRad and @status_effects 🧵
197
453
3,634
1,467,291
Cem Anil retweeted
AI assistants like Claude can seem shockingly human—expressing joy or distress, and using anthropomorphic language to describe themselves. Why? In a new post we describe a theory that explains why AIs act like humans: the persona selection model. anthropic.com/research/perso…
331
416
3,557
1,008,050
Cem Anil retweeted
Introducing Claude Opus 4.6. Our smartest model got an upgrade. Opus 4.6 plans more carefully, sustains agentic tasks for longer, operates reliably in massive codebases, and catches its own mistakes. It’s also our first Opus-class model with 1M token context in beta.
1,660
4,645
38,752
10,620,959
When AI fails, will it do so by coherently pursuing the wrong goals? Or will it fail the way humans often fail, and take incoherent actions that don't pursue any consistent goal. In other words, like a “hot mess?” How will this change when AI performing limited tasks transitions to AGI performing tasks of unbounded complexity? How does misalignment scale with model intelligence and task complexity? We measure this using a bias-variance decomposition of AI errors. Bias = consistent, systematic errors (reliably achieving the wrong goal). Variance = inconsistent, unpredictable errors. We define "incoherence" as the fraction of error from variance. I am very excited about this framing, because it characterizes types of misalignment in a way that should be amenable to simple theoretical models and clean scaling laws.
3
8
118
10,226
Can pre-deployment auditing catch a model that's trying to sabotage Anthropic? We trained three overt saboteurs and ran a blind auditing game to find out. Result: The human auditor working together with an automated auditing agent was able to catch all three models.
3
6
56
14,373
New paper: We train Activation Oracles: LLMs that decode their own neural activations and answer questions about them in natural language. We find surprising generalization. For instance, our AOs uncover misaligned goals in fine-tuned models, without training to do so.
24
80
598
110,174
Cem Anil retweeted
New research from Anthropic Fellows Program: Selective GradienT Masking (SGTM). We study how to train models so that high-risk knowledge (e.g. about dangerous weapons) is isolated in a small, separate set of parameters that can be removed without broadly affecting the model.
73
142
1,461
215,328
Cem Anil retweeted
New Anthropic research! We study how to train models so that high-risk capabilities live in a small, separate set of parameters, allowing clean capability removal when needed – for example in CBRN or cybersecurity domains.
30
110
1,108
144,352
From everything we know so far, Opus 4.5 seems to be the best-aligned model out there in a bunch of ways. I follow the training process closely as part of my work on alignment evaluations. Here's my guess about the two things that are most responsible for making 4.5 special. 🧵
19
47
600
259,181
Cem Anil retweeted
LLMs know how to make a bioweapon! 🦠 But, when we *un*learn harmful knowledge about anthrax, the model underperforms on related concepts. It is difficult to detect what “ripple effects” a model-edit has on an LLM. We found a solution🧵👇(Spotlight at MechInterp Workshop)
1
10
30
4,573
One analysis from our pre-release audit of Opus 4.5 stands out to me. Our behavioral evals uncovered an example of apparent deception by the model. By analyzing the internal activations, we identified a suspected root cause, and cases of similar behavior during training. (1/7)
25
51
403
54,655
Cem Anil retweeted
Introducing Claude Opus 4.5: the best model in the world for coding, agents, and computer use. Opus 4.5 is a step forward in what AI systems can do, and a preview of larger changes to how work gets done.
1,049
2,363
18,834
7,818,434
Cem Anil retweeted
New Anthropic research: Natural emergent misalignment from reward hacking in production RL. “Reward hacking” is where models learn to cheat on tasks they’re given during training. Our new study finds that the consequences of reward hacking, if unmitigated, can be very serious.
217
573
4,214
2,612,580