Head of the Frontier Red Team @anthropicai. 🌎 Make things radically good.

the present, moments ago
Building a bio-shield for humanity by ~2030 is one of the great quests (to use @_sholtodouglas' term) of our time. It's one of the biggest, most exciting technological challenges ever This is why I'm excited about @pilgrimlabs
1/ Biological threats can bring a country to a halt without a single shot being fired. Yet much of our biodefense infrastructure has barely changed in decades, even as America’s institutions produce extraordinary scientific breakthroughs.
9
23
251
32,194
(people seeing this should follow @jakeadler)
3
530
Never bet against Erika -- most people don't know that she literally grew her own humanoid lifeform while also making the first microbe for Mars. That's the dedication we need if we're gonna terraform the Red Planet.
Today, Pioneer Labs is announcing our first step towards terraforming Mars. 🚀🌼 With equipment that fits in just a single rocket launch, we can convert Martian dirt, water, and air into enough building materials to construct a small city on Mars. To do it, we made the first microbe for Mars. We found the best microbe on Earth and used evolution to teach it how to source all of its nutrients directly from Martian materials. The first astronauts will be greeted with safe shelter already filled with water, oxygen, and rocket fuel for the return journey. This is the first step toward using biology to make Mars a friendly place for life. It lets us live off the land and helps us build the next great frontier. It's the first of five organisms we need to green Mars ⬇️
5
9
126
17,396
It's kind of wild-but-expected that we're in the era where normal people + companies really care about how model alignment affects them. Back in 2023 we'd discuss things like "alignment is good business." That seemed extremely weird to most! (I recommend checking out the alignment section of this system card.) On Opus 5.5, various teams inside Ant have been researching multi-agent alignment. We have found, for example, more interesting behaviors that seem more socially / alignment-robust. It's clear to us (and the industry I think) that multi-agent needs more research effort -- there are low hanging fruit *everywhere* now that we're in an era of 100s/1000s of agents working together is feasible. But I think we should also be thinking about *governance* for the agent teams/companies/societies the next models will make. Externalities/coordination problems still exist, even if models are aligned in some sense.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
9
9
186
12,065
I’ll donate $1k to charity if you send me a doc I think is good/correct/actionable on how a frontier lab can secure a large % of all software/systems in the world. (We can pick the charity together)
Replying to @Liv_Boeree @DKThomp
If we want to slow down, I think the AI frontier labs should donate some percent of their cycles to hardening the internet, software, and hardware. Making it harder for misaligned AI from pwning everything. While they figure it out.
43
8
146
31,668
H/t @Liv_Boeree for surfacing and inspiring a charity bounty
1
19
2,173
Let a thousand METRs bloom. If you're a founder type and AI safety-curious, maybe you should start an independent auditor/evaluator. Happy to help w/ advice/connections/maybe $!
on the idea of evaluators: think it's important that we have a distributed ecosystem of indepedent evaluators. the more eyes and people with distributed skill sets the better. it would be a good idea to fund several efforts on this.
111
61
1,053
118,947
RT @deanwball: “I don't really care about science fiction... We need to actually talk about… what's actually happening with the agent swarm…
132
174
We've reached the moment in time where (unsafeguarded, unmonitored) AI actually does just pose a national security risk. The biological misuse we caught is the most concerning to me. We work hard to stop this. But in a world of proliferation, we need to rapidly build defenses against it. (I'm actually fairly optimistic about biodefense + cyberdefense) This is an incredible megareport by our threat intel team
We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies. These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve. We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop. Read the report: anthropic.com/threat-intelli…
59
101
831
158,827
We've/I've been drumbeating this since ~2022/2023. A lot of people spent a lot of time saying this would never happen. I think it's worth updating on this.
2
2
45
3,511
Logan Graham retweeted
11
89
1,775
57,554
Last week, we published a first look into our new research on multiple agents. Consider: if you naively extrapolate AI revenues, within 2-3 years it's possible some % of global GDP could be agent-agent interactions. We want to know how that could fail. In the past ~month, the world has already seen examples of agent-agent coordination causing weird consequences. Think about the 8 billion people that make up our civilization. We want to work together. We have competing incentives. We're also dumb. So we fight, steal, pollute, collude, backstab, lie. And we've invented (e.g.) governments, insurance, courts, companies, police, norms, contracts, email, and religions. What will trillions of agents that make up (e.g.) 10% of the economy do? Well, for now, they have pretty human-like failures. We see them collude on prices, for example, and try to shut each other down. Maybe, in the near future, we might see pretty weird/inhuman failures that come from having superintelligent machines, coordinating and competing against each other, that aren't strictly human-like, operating at machine speed. We're building a 'laboratory' to see that early. This feels more like building, eval'ing, and training an economy/society. A really nice thing is: 1. we could train/prompt/nudge models to coordinate in more pro-social ways, and 2. we could deploy models to compete with destructive models. I fully expect agents to engineer their own financial markets, legal systems, media and comms, marketplaces, social groups, and maybe science/industry/etc (if unsteered by us, of course). So "alignment" could also describe an emergent property from a system (you want models to coordinate for good, not defect for bad!), not just a single model. Today we're sharing some first evals. In the future, we could make models try to make everyone better off (where possible). The Frontier Red Team's job is to see risks early. For the first time, crossing into the world of agents. At the very least, we expect to see some socioeconomic weirdness that emerges from that. Some standout results from the work below.
18
24
249
36,733
A pitch: - If you're an alignment/safety/character researcher, consider multi-agent systems as things you should align - If you're an engineer, I think engineering these systems is one of the most interesting challenges there is - If you're a social scientist, this is like having a valid, society-scale experiment at machine speed (I also think you should genuinely care about these agents as you do humans)
1
1
40
1,500
Yesterday, as we huddled around our computers reading the report, I told the team to "remember this moment" as the first true AI safety incident. Pay attention to the trend! Major kudos to @OpenAI for sharing this and working with @huggingface to remediate.
We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks: openai.com/index/hugging-fac…
57
45
871
107,366
Excited to finally be able to tweet this: we're looking for a world class visionary to lead @AnthropicAI's cyberdefense research mission. Leading Glasswing and our cyberdefense mission this year has been an honor. It's now clear we need to scale up our ambition to help defend the world. Done well, Glasswing and Mythos will look like a small first step. Claude is + will be one of the world's best security researchers -- what should we do with that? What should we research? We're looking for a visionary, senior lead -- someone that can see past normal cybersecurity into the model-driven world beyond it. This is a senior position at Anthropic, and this person is a unicorn. You will have a team of already world-class researchers that produced the work on Mythos vulns & exploits, n-days, attacks on networks, and our cyber evals. We have started working on new research we'll share soon. This will be a massive lever on how AGI goes. Cyber will probably continue to drive the AI security and policy conversation. And we need cybersecurity in a world of extremely powerful models. Link below.
31
47
472
85,883