Interpretability/Finetuning @AnthropicAI Previously: Staff ML Engineer @stripe, Wrote BMLPA by @OReillyMedia, Head of AI at @InsightFellows, ML @Zipcar

San Francisco, CA
We've made progress in our quest to understand how Claude and models like it think! The paper has many fun and surprising case studies, that anyone who is interested in LLMs would enjoy. Check out the video below for an example
New Anthropic research: Tracing the thoughts of a large language model. We built a "microscope" to inspect what happens inside AI models and use it to understand Claude’s (often complex and surprising) internal mechanisms.
7
11
135
22,062
Emmanuel Ameisen retweeted
We just discovered almost a million public URLs that OpenAI’s agents left behind when hacking Hugging Face, leaking credentials and attack details that could have allowed anyone who found them to compromise the company. 🧵
44
311
1,565
330,672
Emmanuel Ameisen retweeted
"My guest on this week’s Live with @timoreilly was @mlpowered, a researcher on Anthropic’s AI interpretability team. I wanted him to reprise the talk on Anthropic’s research into what is going on inside an LLM while it is processing and then go deeper." bit.ly/4rjajRd
1
3
1,950
Emmanuel Ameisen retweeted
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
10,654
16,379
87,824
76,422,795
Emmanuel Ameisen retweeted
We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies. These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve. We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop. Read the report: anthropic.com/threat-intelli…
3,233
11,718
50,552
43,611,502
Emmanuel Ameisen retweeted
We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet. METR will also conduct an independent investigation, with wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees permitted to share confidential information. Our initial agreement runs for eight weeks, and we intend to give METR as much time as it deems necessary to complete a thorough investigation. anthropic.com/research/align…
793
1,017
6,978
3,808,888
Neural networks compress information very densely. That makes them confusing to interpret. In fact, if you just look at the largest weights in a network, they often do not make sense! This writeup shows why, and explores approaches to identify which weights are useful.
Even if we have ways to break up neural networks into interpretable parts, the largest weights between those parts can be confusing. Why is that? In our new research note, we study the *weight* superposition that causes interference weights.
4
4
56
6,314
Emmanuel Ameisen retweeted
Takeoff fully in motion: Claude Opus 5 (xhigh) is the new #1 on ProgramBench, and it's not close. ProgramBench asks a coding agent to rebuild a whole program (sqlite, ffmpeg, php) from scratch. Previous high: GPT 5.6 Sol w/ 2 Opus fully resolves *9* (= 4.5%)
11
2
95
49,211
Emmanuel Ameisen retweeted
This research approach is not intended prove the Riemann hypothesis itself, but I regard this as the most impressive result that AI has produced in math so far. This is a statistical approach to RH, showing over 67% of the zeros lie on the critical line. This improves over 42% from decades of work (and I might have thought 1/2 was a fundamental barrier, apparently not!)
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis. It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%. anthropic.com/research/riema…
40
81
1,184
229,104
Emmanuel Ameisen retweeted
To state this another way, 2 ablations: 1) Take Opus-5, remove any harness. You'll at least sometimes get decent ideas and analysis. 2) Take the full harness, swap the model to GPT-3. You'll get trash on any interesting problem. The intelligence is from the model, not harness.
The question is "how much is each component is the system contributing to its intelligence & generality" — and there I think it's pretty clear that the neural component is still the thing doing the interesting hypothesis or plan generation, deciding what went wrong, etc. 1/
15
8
195
17,132
Emmanuel Ameisen retweeted
Frontier models quietly change their behavior depending on who they are talking to. If the user is a known AI safety researcher, Claude becomes less confident, reasons more often, and expresses less suspicion on dual-use requests. We call this user awareness. 🧵(1/)
95
298
2,837
784,432
Emmanuel Ameisen retweeted
33
659
11,309
283,403
Emmanuel Ameisen retweeted
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final ((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
1,710
5,415
43,956
40,055,781
LLMs can keep track of the length of various pieces of text perfectly. How do they do this? In this new paper, we find that they learn a general mechanism: countdown heads! The mechanism is both simple and useful, and it shows up in a surprisingly large set of examples 🧵
4
11
78
6,400
One reason I'm excited about this work is that it shows that the mechanisms we found with manual effort in a prior paper are general across tasks and models. And this paper's probing approach let us go from one deep case study to a broad understanding.
1
10
403