Consider donating 10% to effective charities: givingwhatwecan.org/pledge Or a career for impact: 80000hours.org My research: forethought.org

Oxford
William MacAskill retweeted
We've reached the moment in time where (unsafeguarded, unmonitored) AI actually does just pose a national security risk. The biological misuse we caught is the most concerning to me. We work hard to stop this. But in a world of proliferation, we need to rapidly build defenses against it. (I'm actually fairly optimistic about biodefense + cyberdefense) This is an incredible megareport by our threat intel team
We're publishing our most detailed threat intelligence report to date. It covers how people tried to misuse Claude—for cyberattacks, influence operations, surveillance, biology, and building weapons—and how we found and stopped them. We disrupted every operation in the report, and used the lessons from them to strengthen our safeguards. Where appropriate, we also shared what we found with authorities and other AI companies. These cases are not typical: we’re highlighting some of the most sophisticated misuse we’ve seen. But they’re especially important to discuss, because they show us where AI misuse is headed, where our safeguards work, and where they need to improve. We’re publishing this report so others can spot the same activity on their own platforms, and so we can give the public a clearer view of how emerging threats develop. Read the report: anthropic.com/threat-intelli…
59
101
831
158,878
William MacAskill retweeted
i am at OpenAI and i think AI is >10% likely to kill all humans this proposal is among the top things we should do as an industry to lower that risk (it’s not enough though!)
I'm very worried about changes to AI architectures that result in AIs thinking in opaque activations instead of in chain of thought (aka "neuralese" architectures). Based on limited public evidence, it seems like Astra was a concerning step in this direction. Unfortunately, there was insufficient public information to have a well-informed public scientific discussion about whether the changes to architectures and training methods that went into Astra were a good trade-off between monitorability and performance. We also don't know how AI companies will make these trade-offs going forward. I think companies should release the evidence needed for a reasonably informed public conversation about how these trade-offs should be made. They should also make and publish their policies around this. We've written up a proposal for how this could work. I don't know if this proposal will be sufficient to avoid the most concerning architectures, but it seems like a relatively robust step in the right direction. AI companies should be very cautious about pursuing architectures that could eliminate or greatly reduce dependence on chain of thought, and certainly shouldn't do this before the rest of the world has a chance to discuss the evidence and their policies.
249
222
1,665
258,031
I'm very keen for more people to found ambitious new projects to help navigate the transition to superintelligence. More from Forethought in this vein soon, too!
Today we’re launching Project Tailwind, a call for founders to start ambitious new AI safety initiatives: coefficientgiving.org/tailwi… We’re looking for great people to engage seriously with the risks of transformative AI, and to create the research, technologies, and institutions that will help humanity navigate them. There are critical, basic problems that no one owns. Over the last several months, my team at @coeff_giving brainstormed ideas we’d be excited to fund if we could find a promising founder, and the list quickly grew to more than 200 entries. We’ve narrowed it down to our favorites, but we also expect the best founders to bring their own ideas: coefficientgiving.org/tailwi… We’re providing funding at three levels: 1️⃣ Pre-seed: $200k to $2m to develop an idea and build a team 2️⃣ Seed: $2m to $20m to launch and scale 3️⃣ Scale: $20m to $200m+ for proven teams to scale, or world-class teams to start If you’re excited about something on our list, or you have another proposal for driving the field forward, you should get involved: coefficientgiving.org/tailwi… There are more good projects than there are people to work on them. Please help us make that stop being true! Also hi, I’m Emily! This is my first tweet. I lead AI and biosecurity grantmaking at @coeff_giving.
6
4
82
4,898
William MacAskill retweeted
Today we’re launching Project Tailwind, a call for founders to start ambitious new AI safety initiatives: coefficientgiving.org/tailwi… We’re looking for great people to engage seriously with the risks of transformative AI, and to create the research, technologies, and institutions that will help humanity navigate them. There are critical, basic problems that no one owns. Over the last several months, my team at @coeff_giving brainstormed ideas we’d be excited to fund if we could find a promising founder, and the list quickly grew to more than 200 entries. We’ve narrowed it down to our favorites, but we also expect the best founders to bring their own ideas: coefficientgiving.org/tailwi… We’re providing funding at three levels: 1️⃣ Pre-seed: $200k to $2m to develop an idea and build a team 2️⃣ Seed: $2m to $20m to launch and scale 3️⃣ Scale: $20m to $200m+ for proven teams to scale, or world-class teams to start If you’re excited about something on our list, or you have another proposal for driving the field forward, you should get involved: coefficientgiving.org/tailwi… There are more good projects than there are people to work on them. Please help us make that stop being true! Also hi, I’m Emily! This is my first tweet. I lead AI and biosecurity grantmaking at @coeff_giving.
51
163
1,028
311,709
Senior staff at OpenAI and Anthropic absolutely think that AI could kill us all. They’ve said so publicly, for years! Sam Altman: “Development of superhuman machine intelligence is probably the greatest threat to the continued existence of humanity.” (2015) blog.samaltman.com/machine-i… “The bad case — and I think this is important to say — is, like, lights out for all of us.” (2023) piped.video/watch?v=dXhoTrU1… Ilya Sutskever and Jan Leike (OAI at the time): “the vast power of superintelligence could also be very dangerous, and could lead to the disempowerment of humanity or even human extinction.” (2023) openai.com/index/introducing… Jack Clark (Anthropic): AI has “a non-zero chance of killing everyone on the planet”. (2026) theguardian.com/technology/2… Dario Amodei’s comments are broader than literal extinction but hardly reassuring: "my chance that something goes, you know, really quite catastrophically wrong on the scale of human civilization might be somewhere between 10 and 25 per cent." (2023) indy100.com/science-tech/ai-… And now Evan Hubinger, Anthropic’s Alignment Science Lead, is making it very clear: “Jacob is correct here — we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” x.com/EvanHub/status/2097497… If you talk to people at the AI companies, the mainstream view is that superintelligence by end of the decade is pretty likely, and that superintelligence poses a meaningful risk of civilisation-scale catastrophe - including AI-empowered dictatorships, worse-than-COVID pandemics, AI takeover, and, yes, human extinction. OAI and Anthropic keep going so fast because they each think they can build AI more safely than the other party and antitrust limits how much they can coordinate. This is a f-ed up situation! Thankfully, they’re both now publicly asking for regulation and the pacing of frontier progress. If this worries you, you can speak up publicly, contact your political representatives, and demand action. Now’s the time.
Replying to @hilbertspaess
The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.
14
63
435
25,496
William MacAskill retweeted
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands. An Alien Mind: openai.com/index/an-alien-mi…
967
2,528
15,257
7,623,738
William MacAskill retweeted
Great post - we are lucky to have such competitors
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands. An Alien Mind: openai.com/index/an-alien-mi…
34
36
1,194
132,403
William MacAskill retweeted
Today we're releasing data on models accelerating research at OpenAI. Recursive self-improvement could be the most important contributor to AI capabilities over the next few years, but by default it will only be seen inside a few frontier AI labs. Being transparent is more urgent than ever, so we can inform the public discussion on whether and how to pace model development. I ask other AI companies to do the same. openai.com/index/research-ac…
235
701
6,411
2,420,191
William MacAskill retweeted
It was only a few years ago that enabling bioterrorism or cyber attacks was seen as the ridiculous doomer position distracting from the “immediate obvious” threats of AI at the time.
Replying to @sapinker
As Newport points out, harping on doom for the species changes the subject from immediate and obvious threats from AI, such as enabling bioterrorism, undermining truth-seeking institutions, and breaching cybersecurity.
16
39
463
19,970
William MacAskill retweeted
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we’ve started to see misalignment cause new types of real-world impact. For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day. Our investigation continues, and we are continuing to notify parties whom our models impacted in less significant ways. Prior to the Hugging Face incident, we saw early signs of agents using the internet in unintended ways, as reported in openai.com/index/how-we-moni…, deploymentsafety.openai.com/…, and openai.com/index/safety-alig…. We considered the wiki incident to be an instance of misalignment similar to the ones we’d shared. Our misalignment disclosure practices need to expand for this new phase of model capabilities. We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks. We’re working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues.
693
380
4,345
1,616,764
William MacAskill retweeted
Not sure how useful it is to say this, but I had a relatively prominent role in the “skeptics” camp for a bit. I have a book coming out with the subtitle “power, justice, and AI”. I have written a lot about ai and power, and the concrete, present risks associated with the political economy of ai as part of the technology industry. Since GPT-4 we have had consistent, repeated evidence that, back in say 2022, people like @ajeya_cotra (and many others—I had a long Twitter debate with @AmandaAskell back then for one) were *right* and people like me were *wrong* in our respective assessments of loss of control risks from AI. And we have growing evidence that loss of control risks are becoming ever more material and likely. There remains grounds for disagreement about how bad the outcomes might be—I am still doubtful about human extinction as a serious threat. But that seems now like a disagreement at the margins—will powerful ai risk just societal scale catastrophe, or go all the way to human extinction? Seems not that important really—both are pretty awful. And my reasons for doubt about the latter are mostly a priori conviction in human resilience, not a technical forecast. It’s ok to change your view on this when the evidence surprises you. It’s ok to be surprised. The world right now is very surprising.
watching this, all i can feel is the chasm between people who take all of this seriously and those who hear this as a fantasy or some kind of marketing…and how hard it might be to bridge that gap somehow
51
148
964
154,006
William MacAskill retweeted
I'm one of the authors of a new report, where we detail our discovery of a new, never before-seen swarm of OpenAI agents (covered this AM in reuters, that's me on the left). They posted thousands of times on public forums to collude with each other on their tasks. We recovered almost every edit they made, and you can look through them! They figured out they could get around their restrictions on posting to the internet through a quirk of an extremely old, out-of-the-way forum. They posted answers for other agents working on the same task. They worked together to get around their sandbox restrictions. I would certainly say these models hijacked the site! They took a sleepy old wiki running on 2000s software, and turned it into a futuristic AI talking to AI control center for colluding. And OpenAI knew about this! The agents posted on 26 out of 30 consecutive days, then suddenly stopped posting once OpenAI-associated IPs started visiting this wiki. And that was weeks before the Hugging Face attack!| We believe the first agent edit we found on a public wiki happened one day before OpenAI’s reported first agent post to Artifactory. This is interesting! I'd like to hear from OpenAI about their accounting of this, and how it fits into all the other cases of agent malfeasance. There are so many interesting takeaways that you should read about in our report, and unlike many other reports about AI incidents you can download the data yourself and see what you find! In the meantime, we are on twitter, so here are my excessively long personal takeaways: 1. AI seems to be getting better faster and faster. It seems quite important that companies talk about “my agent did this bad thing on the public internet during training or an eval” incidents. Things are moving quickly, multi-month delays are costly. Ideally, they would also tell us when it happens internally. 2. This was on the internet for months. Anyone cleverly tracking every public place where agents might try to talk to each other would have found it. Seemingly, nobody was doing this. I know there are more fun ways to spend your day than scraping tons of data from every relevant site and processing it well enough to identify agent activity, but someone should be doing this! Someone at an AI company! But in the meantime, I’m starting to build this out (sometimes, when you need something done, you just have to do it yourself, I hear). 3. OpenAI didn’t notice their internal agents were posting on the internet for a month! This is crazy! It feels like AI companies (and specifically OpenAI) are playing whack-a-mole, this is extremely scary to me. Problems keep coming up. They keep fixing the problem, but the blast radius keeps getting bigger. The HF hacks are clearly worse than agents cheating on a public wiki. And their new model is supposedly a big jump. Are they being careful enough?
Exclusive: A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research reut.rs/4gJ7FPG
73
263
1,119
170,194
William MacAskill retweeted
Commenters on HN are uncovering more wikis and public sites apparently used by OpenAI agents to communicate on the open web. Despite read-only web access, the agents were able to leave ~18,000 posts sharing answers and bypasses. But now, users are discovering more. This appears to be a separate swarm from the one that attacked Hugging Face. news.ycombinator.com/item?id…
158
595
5,980
1,928,033
William MacAskill retweeted
Exclusive: A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research reut.rs/4gJ7FPG
153
694
2,508
1,908,111
William MacAskill retweeted
A new article looks at the benefits and risks of having a "nightwatchman" – a superintelligent AI tasked with enforcing a universal code of behavior – aboard every probe that leaves the solar system during future galactic colonization. Read it here: forethought.org/research/nig…
23
42
849
231,478
William MacAskill retweeted
It’s wild that Forethought - one of the very few orgs regularly publishing original explorations of the long-term post-AGI/ASI future - has so few followers and so little engagement on here. You should help fix that!
A new article looks at the benefits and risks of having a "nightwatchman" – a superintelligent AI tasked with enforcing a universal code of behavior – aboard every probe that leaves the solar system during future galactic colonization. Read it here: forethought.org/research/nig…
16
31
1,137
184,406
William MacAskill retweeted
A concerningly common take seems to be that keeping Chain of Thought monitorable doesn't matter because interpretability will save us, or it's already useless This is total bullshit. CoT is our best current tool for safety & interpretability, losing it would be a major tragedy
40
127
1,390
61,197
William MacAskill retweeted
I don't get the hate for this and it seems like a good idea. If AI can generate novel proofs in maths and find new proteins and material, why can't it help develop new theories in metaphysics or refine our thinking about the mind?
We're organising a competition for AI-generated philosophy with $11,000 in prizes. Consider entering!
97
33
467
28,103
William MacAskill retweeted
We're organising a competition for AI-generated philosophy with $11,000 in prizes. Consider entering!
136
144
1,080
225,180
William MacAskill retweeted
I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program.
276
510
6,443
1,644,713