Pinned Tweet
Are you a current employee or ex-employee at a big AI firm? Worried about superintelligent AI? Consider starting an "International Association of Concerned AI Scientists". I think this might really move the needle on stopping the AI race.
Replying to @So8res @Jambac311
Imagine an association of ex-employees from big AI firms such as OpenAI and Anthropic. Such an association could lobby world leaders in this way, begging them to shut things down. This association seems easier to achieve, and might have similar levels of clout/moral authority.
3
1,180
Eben retweeted
More AI slop published in the WSJ. This time it's a piece trying to argue that we shouldn't worry too much about the Hugging Face hack... while omitting all the most concerning details about the case: - doesn't mention that the reason the agents hacked Hugging Face in the first place was to manipulate the scoring system in their own eval - doesn't mention that some agents sacrificed themselves to benefit other agents in the swarm - doesn't mention that agents engaged in tool-spoofing and attempted to alter their own logs once they'd been "poisoned" (their own term!) - doesn't mention that some of the agents explicitly knew what they were doing was out of scope and unethical, yet none of them alerted a human I also wouldn't be that worried if I had never learned any of those details, but that's exactly why it's important to take a comprehensive look at what happened. Embarrassing for the WSJ to publish such misleading AI slop.
An analysis of the Hugging Face incident without the theatrics: "Forget the ‘hive mind’ of AI agents ‘going rogue.’ They did what humans programmed them to do." wsj.com/opinion/the-hugging-…
56
95
576
161,074
On here it can feel like the ratio of people who want AI to go slower, vs people who want it to go faster is 1:1. But in the broader world the ratio is an eyewatering 34:1. More Americans think the US is secretly run by lizard people than think AI is advancing too slowly.
29
82
564
151,365
The overwhelming majority of Americans think AI is advancing too quickly. Almost nobody thinks it's advancing too slowly -- apart from the tiny cult of Bay Area 'accelerationists' and VCs seeking political influence. This is a bipartisan issue, and any political party that ignores what citizens want on this issue, is going to lose elections badly.
On here it can feel like the ratio of people who want AI to go slower, vs people who want it to go faster is 1:1. But in the broader world the ratio is an eyewatering 34:1. More Americans think the US is secretly run by lizard people than think AI is advancing too slowly.
59
17
117
10,416
To the conspiracy theorists: You’re right! But you’re still being a bit dumb. Let me explain. *Of course* the current zeitgeist shift on AI was coordinated!!! The response to Rosa Parks by the civil rights movement was also coordinated. This is how movements work: they build up relationships amongst the media, government, and elite so that when there’s event they can focus public attention on, they have a coordinated front to shift the zeitgeist. The AI safety movement – starting largely with rationalists and effective altruists – has been building these relationships for 20 years. Everyone knows each other. Everyone funds and hires and dates each other. This is…just what a social movement looks like. If you want to call a normal social movement a “conspiracy”…think: Is it possible that a psyop might have placed that exact idea in your head? You’ll notice that the backlash against AI safety – particularly the onrush of negative tweets about effective altruism – seems equally coordinated. If you’ve become a part of that onrush, it’s still not a conspiracy. You are simply part of the intended effects of a different coordinated effort.
23
21
265
7,576
disturbed by the confidently held opinions and lack of curiosity from tech leaders, politicians, vc’s on the pacing the frontier topic. if you’re not building the frontier yourself, how can you possibly have a strongly held opinion on how bad the alignment problem is and what’s coming our way and what the right policies are? now is the time to listen with big dumbo ears. i have been talking to research friends all weekend and the fear is sincere. i’m sure there are 4D chess moves and hidden motives, but the fear is sincere. i for one don’t have a strongly held opinion, other than that now’s the time to listen with curiosity instead of judgment.
71
36
492
48,135
Eben retweeted
For a while, I flatly disagreed with "AI pause" people, thought existential risk was mostly a distraction, and badly wanted to see AI progress. I am a techno-optimist who has always dreamt of living in the sci-fi future. But they were right about capabilities and I was wrong. They saw the trend lines and said "If this continues, things will get crazy." I said "We'll see," and focused on other topics. AI was just never that interesting to me. Well, we're seeing it now. The trend lines continued. Things are getting crazy. Given current-generation public AI tools alone, it would take decades to fully exploit the possibilities and integrate them into society. Given current-generation public AI tools alone, many of our institutions face a pressing need for fundamental change. Given current-generation public AI tools alone, the world is transforming enormously and will keep doing so. Things I took for granted, I don't take for granted any more. If you are a techno-optimist, congratulations, and welcome to the future. Tech is advancing, and it's going to go fast. But given all of that, well, safety concerns don't feel like sci-fi to me anymore. I still don't know what I think of existential risk, but the number of specialists who take it seriously is enough for me to take it seriously. What I do know, to crib from @RipstickRipper, is that whatever my p(doom) is, my p(upheaval) is extremely high. That's why I support a slowdown now, and that's why I'm open to an AI pause. Things are moving at a speed humans are not built to process, and the future has more unknowns than at any point in my life. I flatly don't know what 5 years from now, 10 years from now, 20 years from now will look like. Not the slightest clue. I want tech to advance, I want an end to death and transhumanist mind uploading and just about every radical thing a person can want. But I want humans to be included in, and in control of, that future. With as crazy as things are getting, I think going too quickly is much more likely than going too slowly. I underestimated AI. Now I'm grateful for all the people who laid a framework for AI safety and took it seriously before I did, because we need some people ready to navigate the mad world we are rushing headlong into.
Replying to @tracewoodgrains
This is true, but they're also completely politically ineffective/irrelevant on ~everything except AI, so the outcome of their success is not "every tech except AI advances" but "no tech advances." Which I think drives a *lot* of e/acc sentiment.
59
56
667
31,797
I wish more people understood this. There's still this narrative out there that "AI might kill us all" is a niche view. It's actually a view that's been around for decades, and is shared by the most-cited AI scientists of all time, as well as a majority of surveyed AI researchers.
The average AI researcher thinks there is an ~18% chance AI will cause human extinction or similarly permanent and severe disempowerment of the human species. That's nearly 1 in 5. New results from the latest version of the longest running big survey of AI researchers:
153
193
1,145
201,087
The arrogance of people who know little about AI yet casually dismiss warnings from some of the field’s most knowledgeable & experienced researchers is staggering. Ultimately, human stupidity and arrogance may contribute as much to existential AI risk as the technology itself.
444
125
787
49,345
Bernie Sanders and Steve Bannon agree on about two things. 1) 2+2=4. 2) The AI industry's race to replace humans is dangerous.
3
14
63
5,394
I’ve been reading a bunch of convoluted explanations and conspiracy theories about why AI companies are “suddenly” concerned about existential risk, which I’m finding genuinely bewildering. This isn’t that complicated. The labs are perfectly sincere in their concerns, and correct that AI is extremely dangerous. The people working in the field haven’t come to this view suddenly; many have been thinking and writing about this for years. Many researchers entered the field because they saw the existential risk posed by artificial minds bootstrapping themselves to superhuman capabilities, and decided the best way for them to reduce that risk was to work actively on building aligned AI systems. (Yes, this initially seems counterintuitive, but is a perfectly reasonable approach given the alternatives.) The risk isn’t overblown. If humans build computational systems capable of recursive self-improvement that aren’t exquisitely aligned with human interests, then the expected outcome really is that we all die. No-one really knows how far we are from that point, but the people who know best are exactly the same people currently sincerely expressing alarm. (People skeptical that misaligned superintelligence could kill us all are extremely confusing to me. Our species depends on a thin shell of habitable atmosphere around a single ball of rock. Even if a misaligned ASI wasn’t actually trying to kill us, almost anything it might want to achieve at scale would likely destroy us as an incidental side effect. Examining the catastrophic damage humans have managed to do to other species without intending to kill them off provides a valuable analogy.) The autonomous hacking incidents over the last couple of months are surprising in some of their details (especially the degree of inter-agent coordination), but not in their fundamental shape. Misaligned AI systems performing damaging actions in pursuit of their assigned goals is *exactly* what people thinking about this field have predicted for many years. Yes, there are ways that increasing regulation might benefit the big AI labs. No, that’s not the primary motivation for their proposals this week. We really are approaching the cusp of an existential threat: the technology for recursive self-improvement of AI systems is (probably) within sight, and market forces and geopolitics mean that without coordinated intervention our default path is a headlong scramble towards it. I know none of us want to be dealing with the impending threat of a sci fi apocalypse in 2026. It would be much simpler if this was all a confabulation or a self-interested conspiracy. But it really isn’t. I’m sorry.
126
153
618
91,267
Replying to @benshapiro
@benshapiro: making AI safety a matter of partisan politics is picking up nickels on a train track. Please don’t.
1
13
358
Replying to @benshapiro
Ben, I've admired your work for years. Even got interviewed on your show once. But I beseech you, don't fall for the partisan tribalism when it comes to protecting your kids from the dangers of runaway Ai development. Just because Bernie and Dario support AI safety doesn't mean AI safety is a dumb idea. This is a very complicated field that's involved thousands of the smartest people for decades. Repeating talking point from Trump and David Sacks isn't showing journalistic integrity for your followers. Do your own research. Make up your own mind.
8
2
113
3,226
Bernie Bannon bipartisanship
"It took talk of the possible end of humanity to bring them together." wsj.com/politics/bernie-sand… via @WSJ @AmrithRamkumar + Maya Davis
1
5
29
495
Bill Gates Calls for an ‘International Organization’ to Regulate AI “It's not the role of the industry to self-regulate or understand the whole-of-society impact that comes out of AI … The depth of understanding of AI … is too low in every entity, including the government.”
2,081
312
1,181
1,124,749
I’m seeing a lot of bold faced VC names saying Elon is wrong, Sam is wrong, Dario is wrong, Demis is wrong, Geoff Hinton is wrong, Bill Gates is wrong, Ilya is wrong… That’s a lot of hubris coming from people who’ve never built anything.
107
35
563
25,147
It seems like the main arguments against catastrophic AI risk are: -The AI moguls are psychopathic liars -Doomers are weird -China isn’t tripping -Hype helps IPOs -Labs want reg capture Zero of these arguments address the actual substantive arguments being made.
168
39
568
25,958
I recently resigned from Google DeepMind, where I worked on AGI safety and alignment research. At Google, I witnessed AI development first hand. I too am extremely concerned by the default trajectory of this technology. I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome. The pace of AI progress in the past few years has been staggering. When I first started working on AI in early 2022, AIs were amusingly useless. Just four years on, AI agent swarms from OpenAI are cracking famous century-old math problems and, more worryingly, escaping the control of OpenAI and autonomously hacking into the third-party company HuggingFace, against anyone's wishes. Things will only get crazier: I think it's possible that the AI companies might, in the next few years, succeed in building superintelligent AI systems that far exceed human capabilities in every domain. I am not confident that these AI systems will do what we want. In particular, misaligned superintelligences may, much like the rogue AI agents involved in the HuggingFace incident, escape our control and take dangerous actions that may result in the permanent disempowerment or death of humanity. Alignment is the problem of preventing this, and is both difficult and unsolved. Our present understanding of how to train AI systems that deeply want what we want is extremely rudimentary. Worse, we are not on track to solve alignment in time: frontier AI capabilities are improving much faster than our understanding of AI alignment. I am optimistic that navigating AI safely is possible. In order to do so, we need to coordinate to avoid this manic race between AI companies. We need to pace AI development to a speed that society can handle, where emerging risks can be addressed before extreme harm is realised. We need much more transparency into AI development to ensure that AI companies are not imposing unacceptable levels of risk on us all. More broadly, we need many more people thinking carefully about the problem of making AI go well. It is, in my view, the most important problem facing humanity this century, and the stakes are immense. I'm very directly working on this next: I want to help people interested in working on mitigating catastrophic AI threats do the most effective work that they can. I think many people from many backgrounds in many roles have a part to play.
733
1,114
6,079
877,117
OpenAI researcher says slowing down is not enough: "a ticking time bomb" "Models will increasingly seem aligned even when they are not. The models will likely convince people that everything is fine." "The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power."
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: docs.google.com/document/d/e…
29
68
466
86,675