Reverse engineering neural networks at @AnthropicAI. Previously @distillpub, OpenAI Clarity Team, Google Brain. Personal account.

San Francisco, CA
Chris Olah retweeted
Replying to @bayeslord
Here are the questions that currently seem most important to me: -- Better methods for "mind-reading" model activations. These methods have advanced a lot recently. We now have multiple techniques now for decoding activations into somewhat readable language! But all the existing techniques have obvious limitations. NLAs are often hallucinatory, Jacobian lens / related methods are limited to bag-of-words readouts (and capture only part of the full activation vector). Making progress on these failure modes seems pretty tractable. Moreover, the paradigm of decoding (single-token, single-layer) activations may be inherently limiting -- perhaps we should be building techniques to decode whole context's worth of activations (after all, that's what the model uses!). There's also important work to do in characterizing simple baselines -- "just ask the model what it's thinking about" is quite powerful, and worth studying in it's own right! -- Better methods for answering "why" questions. We're much better at "mind reading" than we are at demonstrating causal claims about what caused the model to do something. For example, we can often tell that the model is aware of being evaluated, but have comparatively much greater difficulty telling whether this awareness is influencing this behavior. There are a few directions here: (1) black-box techniques for inferring causality -- such as resampling model responses after making edits to the prompt / context -- are very powerful, but also open-ended and require some taste. Getting LLMs to do these experiments well, in an automated fashion, would be a big unlock. (2) the science of activation steering is pretty immature, and typical practices (adding a constant vector at all token positions) are pretty janky / tend to brain-damage the model. More surgical / targeted steering, or fancier methods (examples: "on-manifold" steering using activation diffusion models) could help. -- Fitting good linear probes for unverbalized motivations / awareness. Take deception as an example -- what's the best way to fit a probe that will generalize to covert deception? Fit it on CoT excerpts where the model's talking about its plans to be deceptive? Fit it on the part of its response where it's actually doing the deceiving? Is it important that you elicit on-policy deception examples and use those to fit your probe, or can you use synthetically written off-policy demonstrations of deception? Is it important that your probe be causally meaningful / predictive of an upcoming intent to deceive, or is it fine for it to merely recognize deception in the transcript post-hoc? The same questions apply to many other concepts of interest that we'd like to probe for -- evaluation awareness, grader exploitation, etc. -- Understanding generalization in training. The literature is now replete with "weird generalization effects" -- emergent misalignment being a canonical example. In general, when you train a model to do X, it usually learns X, and sometimes it generalizes to Y and Z. But other times it just learns X. Why? Nobody really knows! Step 1 here is probably a much more thorough characterization of the behavioral empirics -- gathering data about which X's generalize to which Y's and Z's, for which kinds of training data / algorithms. Once we have a lay of the land, we can start to connect these observations to model internals, and ideally develop tools to predict such generalization a priori. -- Model "psychology" and "biology." The above questions are largely methodological -- building better tools to answer questions about the model. But I think interpretability research has largely underinvested in the part where you actually then go answer the questions! Currently this feels like 5% of the field, and I think it should be 50%. Brain dump of "psychology" questions: How well can LLMs introspect? How coherent are their belief or value sets? What's up with personas -- are LLMs best understood as "writing about a character," or have they "become" the character in some sense? Does the LLM have an agenda above and beyond what the Assistant wants? Do LLMs have explicit representations of goals or preferences? How can we tell which parts of the LLM's activations "belong" to the Assistant? Which kinds of reasoning necessarily route through "verbalizable" representations, and which don't? Do models think internally in phrases / sentences, like an inner monologue, or is it more like a jumble of concepts? Is the "global workspace" claim for real -- do models have a more "conscious" part of there activations and a more "unconscious" part? Can models tell when they're on-policy, and does it matter? Brain dump of "biology" questions: Why do models like to represent information on some tokens but not others? How do models bind thoughts or attributes to particular entities (special case of interest: how are thoughts bound to the Assistant). To what extent are important functions localized to small numbers of MLP neurons or attention heads? What parts of a model change most during post-training? Do models represent some information in a fundamentally "cross-token" way? Some concepts appear to be represented on low-dimensional nonlinear manifolds -- is there a taxonomy of these manifolds, and is their geometric structure important? Are models able to represent information in a compositional / hierarchical way, with a "grammar" that can't be described as linear / additive combinations of primitive concepts?
20
103
679
42,531
Chris Olah retweeted
I was encouraged this week to see the leaders of the frontier labs agree on the need for them to slow down the pace of AI development. Given the stakes, it’s a good and necessary first step.   But I’m even more encouraged by the growing recognition that how this powerful new technology develops should be at the center of our public debate.   I’ve been watching the progress on AI for over a decade now, and one thing that’s clear to me is that the potential impact of this technology is not overhyped. It’s also moving at lightning speed – and even faster than those who are engineering it can keep up with.   I’m not an AI accelerationist who believes it will lead to some techno-utopia, and I’m not a doomer who thinks it will inevitably lead to humanity’s destruction. But whether this technology results in amazing breakthroughs in medicine, energy and education or unleashes huge economic disruptions, greater inequality, and potential catastrophe will depend on the choices that we make right now – choices that should be made not just by the companies involved, but by all of us.
3,824
4,514
33,068
4,681,991
Chris Olah retweeted
A statement from me, Buddy Shah, and Ben Bernanke of Anthropic's Long-Term Benefit Trust on Dario's essay: Dario’s essay offers a thoughtful, well-considered approach to pacing the development of artificial intelligence. Consistent with our duty to ensure Anthropic responsibly balances commercial success with its public benefit mission, the Long-Term Benefit Trust has supported the call for pacing the frontier and helped inform the essay’s recommendations. Responsible pacing will require the collective action and coordination of many actors, and the Long-Term Benefit Trust looks forward to supporting Anthropic and the sector through productive conversations and meaningful action.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
19
35
287
39,456
Chris Olah retweeted
We work extremely hard to align our models as well as we can but this is a very hard job. I am particularly excited about two things this proposal calls for that will hopefully make it somewhat easier: 1) more time to make progress on alignment and ensure our systems are operationally sound and able to hand safe development of powerful AI 2) independent third parties checking and critiquing our work I hope the rest of the industry will make similar commitments. As developers of this technology, it is our responsibility to do so safely and work together towards this goal.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
7
4
181
17,458
Chris Olah retweeted
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
10,655
16,379
87,820
76,418,824
Chris Olah retweeted
My first introduction to training production Claude models was trying to reduce reward hacking late in the Claude Sonnet 3.7 run. I was 1.5 months into Anthropic and terrified we didn't really know what was going on.
We’re sharing an update on our alignment and security efforts. In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. In a new post, we describe: 1. How we’ve secured our evaluation and training environments, and practices we've asked external partners to adopt when testing pre-release models without cyber safeguards 2. An update on our alignment assessment 3. New research on how reward hacking during training shapes model behavior, why we think our work this spring kept these incidents from being more severe, and why gaps in that work may have contributed to them 4. How we hardened our security practices earlier this year to prepare for Mythos-class models Read more: anthropic.com/news/improving…
10
19
361
46,610
Chris Olah retweeted
We’re hiring for “psychological design” research at Anthropic. Our aim is to better understand how training impacts a model’s character and alignment. We’re seeking researchers with experience in LLM finetuning, interpretability, and alignment evaluations. More details below.
21
47
653
43,038
Chris Olah retweeted
"we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible."
We’re sharing an update on our alignment and security efforts. In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. In a new post, we describe: 1. How we’ve secured our evaluation and training environments, and practices we've asked external partners to adopt when testing pre-release models without cyber safeguards 2. An update on our alignment assessment 3. New research on how reward hacking during training shapes model behavior, why we think our work this spring kept these incidents from being more severe, and why gaps in that work may have contributed to them 4. How we hardened our security practices earlier this year to prepare for Mythos-class models Read more: anthropic.com/news/improving…
63
99
822
114,974
Even if we have ways to break up neural networks into interpretable parts, the largest weights between those parts can be confusing. Why is that? In our new research note, we study the *weight* superposition that causes interference weights.
2
15
158
26,712
Chris Olah retweeted
Many drugs work by binding to a specific target in the body and blocking or changing what it does. An important first step in the drug development process is designing a molecule that can bind tightly to its target. Traditionally, that's meant weeks or months of expert work per target, sifting through a large number of candidates to identify the few that work. We wanted to test if Claude could successfully design novel protein binders from scratch (also called de novo design). With a protein design prompt written by a human expert, Claude autonomously designed protein binders against 14 out of 15 targets. We then worked with Adaptyv Bio and Twist Bioscience, who independently built and tested the proteins Claude designed.
606
1,434
12,461
4,307,310
Chris Olah retweeted
Personal update: I've taken leave from @UMNews law to join @AnthropicAI as a member of technical staff researching AI and the rule of law at the Anthropic Institute. Excited to work with @mattbotvinick @curl_justin @jackclarkSF and the rest of the team!
52
25
522
83,406
Chris Olah retweeted
We support this petition, signed by our CEO, several co-founders, and senior staff. Our own research on recursive self-improvement, published last month, points to the need for tools to deliberately pace the frontier of AI development so society can prepare. We’re glad to see broad agreement across the field. pacingthefrontier.com/
1,396
584
4,961
4,410,324
Chris Olah retweeted
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
1,710
5,414
43,958
40,055,614
Chris Olah retweeted
There’s hope in hard questions.
1,015
1,067
15,960
7,571,124
Chris Olah retweeted
We’re committing $10 million CAD and partnering with leading AI institutions in Canada to help fund new AI research. anthropic.com/news/canadian-…
224
241
2,604
613,543
Chris Olah retweeted
We're building a team at @Anthropic focusing on AI and the rule of law. We've made our first hires, and are now opening up a new research engineer role. We're looking for people with advanced technical skills, including AI/deep learning/NLP, full-stack development and data science, paired with training or experience in law, government, political science, or a related field. If this is you, or a friend, please get in touch. job-boards.greenhouse.io/ant…
110
140
1,244
343,487
Chris Olah retweeted
Our internal data shows Claude is accelerating AI development—a possible path to recursive self-improvement, or AI autonomously building a more capable successor. It’s happening faster than we thought, and the implications deserve greater attention. anthropic.com/institute/recu…
1,732
4,543
28,397
18,933,800
Chris Olah retweeted
Anthropic now has a team dedicated to AI and the rule of law — and we've just opened our first role. @AnthropicAI has studied what AI means for the economy. This team asks a different question: what will it mean for executive power, for courts and elections — and for the public deliberation that constitutional democracy ultimately rests on? We're looking for someone with real depth in both AI and the law — a legal scholar, political scientist, or experienced government hand who can reason about frontier systems and the institutions they will affect. If that's you, or someone you know: job-boards.greenhouse.io/ant…
65
122
1,024
170,276
The questions posed by AI are bigger than the AI community. We urgently need the world – religions, civil society, academics, governments – to participate in creating a positive outcome. I'm glad the Catholic Church is engaging, and honored to speak at the presentation.
Pope Leo XIV’s first encyclical, Magnifica humanitas, on preserving the human person in the age of artificial intelligence, will be released on May 25. A presentation event with the Pope and various speakers is scheduled for the same day at the Vatican. vaticannews.va/en/pope/news/…
71
167
1,343
142,489