Associate Professor at Stanford. Quantum Information and Quantum Gravity. Currently at @Openai

In 1993 the Superconducting Supercollider was cancelled. Estimated cost: $8 billion. An exodus of physicists left to Wall Street, bringing fancy maths and dubious risk management. 15 years later the global financial crisis cost ~$20 trillion. This is why you don't defund physics!
72
484
2,554
Geoff Penington retweeted
OpenAI researcher @boazbaraktcs on what human mathematicians do in a world of mathematical superintelligence: "It will be kind of pointless if AI proves theorems that only AI can understand. So we want AIs to explain to humans." "It will be easier for AIs to explain their work to expert mathematicians than to the general public." "So these mathematicians will be the ones that first interact with AI and explore what are the right questions to ask, and then they, together with the AIs, explain it." "I would still be more interested in hearing Terry Tao explain his own perspective on math results even if GPT-7 could explain it just as well. No matter how good AI gets at writing, I would probably be more interested in reading a book written by a human." @OpenAI
5
3
55
12,522
Geoff Penington retweeted
Anthropic's implied valuation in secondary markets fell around 5% immediately after Dario published We Must Pace the Frontier (onchaintimes.com/pre-ipo-per…). This seems like important evidence (albeit weak, secondary markets are illiquid, weird and opaque) that the market does not think Pacing the Frontier is in Dario's financial interests. I think it is clear that Dario/Sam/Elon are not financially incentivised to call to slow down development. They are calling for slower progress because they are genuinely scared of AI's risks, not because of regulatory capture. Interestingly, I don't think Andreesen and Sacks have a clear financial interest in their position either*. If the frontier gets regulated, AI could become more commoditised, which might be good news for the application layer companies and neolabs that form such a large part of a16z's portfolio. Both sides of this debate seem ideologically motivated rather than financially motivated to me. People should spend more time evaluating AI doom arguments and less time scrutinising financial incentives. * OTOH I think Jensen does have a clear financial incentive to be anti slowing down, and this does seem important to me to highlight.
Replying to @sriramk @micsolana
I mostly agree (mostly you should engage with people's arguments not their incentives) but I think it's not completely symmetric. There's a long history of industries downplaying risks to suit their financial interests (tobacco, oil companies) and no examples that I know of other than AI of industries overstating risks against their financial interests. To put it differently, it's much more established that financial interests can corrupt people's thinking/statements than other interests. Obviously, people are claiming that Dario/Sam/Elon are financially motivated (they're trying for regulatory capture) but I think the view that their position would enrich them is much weaker than the case that Jensen's position would enrich him, which is where the asymmetry lies. In fact, since they came out in favour of pacing the frontier, I have heard that the valuations of Anthropic and OpenAI have fallen substantially in the secondary markets!
4
25
248
22,894
You learn a lot about politicians from the quality of their opinions on “new issues” where there is no established party line from them to follow. It’s not surprising to me that Obama is way ahead of any current politician on AI policy but it is very noticeable
Obama on recursive self improvement and the risks associated with it. When he talks about the urgency for us all to have a say in how this technology evolves and how it is overseen, this is part of the reason why: “The models are going to get smarter and smarter and better and better at a much faster pace, at an exponential pace.   What are the risks of that? There’s the big science fiction risk, wow, these models get smarter than us, and they decide humans are fine, but not necessary. They start setting their own goals, and the killer robots kill us, or we bow down to them.   I do not want to exaggerate that particular risk, but I will say there is a non-zero risk of non-zero chance of that happening, but that’s not actually the risk that I’m most concerned about, although it’s the risk that gets most attention.   And the reason that’s a risk is not because the computer models are conscious, necessarily. It doesn’t mean that they are necessarily feel malice towards humans. It’s just that if they start setting their own agendas, you may get a misalignment between what they want to do and what we want them to do, and that gap can be dangerous. That’s problem number one.   The more serious problem is these models are getting powerful enough that if they get in the hands of bad humans, they can do bad things. They can be weaponized in certain ways. They can do a lot of mischief.”
62
75
1,362
64,440
Geoff Penington retweeted
I've written a blog post responding to the letter about maths and AI signed by 25 Fields medallists. As with the Leiden Declaration, I didn't sign it, but I agree with much of it and welcome its existence. gowers.wordpress.com/2026/09…
37
163
869
233,252
Geoff Penington retweeted
I have signed a letter from 42 fellows and foreign members of the Royal Society to Paul Nurse, the president of the society, concerning the need to treat the serious risks of AI as an emergency. docs.google.com/document/d/1…
56
147
871
257,668
Geoff Penington retweeted
As the text makes clear, contrary to the headline this “sale” is in fact a gift.
16
64
781
42,552
Dan Selsam is one of the most brilliant AI researchers I have ever met. Just a few months ago he was very sceptical about model capability trajectories, and felt that radically new paradigms were needed for progress. You can disagree with him but everyone needs to read this
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: docs.google.com/document/d/e…
8
17
264
17,293
Geoff Penington retweeted
There's no such thing as a winnable war It's a lie we don't believe anymore Mister Reagan says, "We will protect you" I don't subscribe to this point of view Believe me when I say to you I hope the Russians love their children too
9
6
138
11,218
Geoff Penington retweeted
I was encouraged this week to see the leaders of the frontier labs agree on the need for them to slow down the pace of AI development. Given the stakes, it’s a good and necessary first step.   But I’m even more encouraged by the growing recognition that how this powerful new technology develops should be at the center of our public debate.   I’ve been watching the progress on AI for over a decade now, and one thing that’s clear to me is that the potential impact of this technology is not overhyped. It’s also moving at lightning speed – and even faster than those who are engineering it can keep up with.   I’m not an AI accelerationist who believes it will lead to some techno-utopia, and I’m not a doomer who thinks it will inevitably lead to humanity’s destruction. But whether this technology results in amazing breakthroughs in medicine, energy and education or unleashes huge economic disruptions, greater inequality, and potential catastrophe will depend on the choices that we make right now – choices that should be made not just by the companies involved, but by all of us.
3,823
4,514
33,069
4,681,128
Geoff Penington retweeted
There are two ways AI progress could go very badly and that we must avoid. First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must always serve people. To ensure that, we need ways to ensure that alignment and safety techniques stay ahead of progress in model capabilities. Second, we could end up in a world with too much concentration of power. If an extraordinarily powerful AI is used by one person or company to impress their worldview onto everyone else, the results could be extremely dystopian. Avoiding these two threats requires walking a narrow middle path; for example, one country could gain too much power. Another example is one lab ending up with too much power.
The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to work if we cannot. We welcome a federal framework that sets consistent safety requirements for frontier AI. But we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence. Consistent rules to manage frontier risk so that we can maximize the benefits are a good idea (and we are excited by ideas like independent auditors). Years ago, companies like ours developed things like Responsible Scaling Policies and Preparedness Frameworks. Those were good for that moment, and focused primarily on the deployment of completed models, not what happens during their development process. Today's shift to focusing on safe development and evaluation will need new tools. For example, at OpenAI we now formulate explicit safety cases in advance of frontier reinforcement learning runs we expect to significantly increase capability, in addition to the safety work we have long done in advance of model releases. We hope that other companies will learn from our approaches and propose their own; we think shared standards for misalignment, monitoring, and safety will lead to better outcomes. We look forward to collaborating with our colleagues across the industry to formulate the best version of these. When we talk about “pacing”, we do not mean “stopping”. Progress has been rapid and will continue to be. But it should be slower than it otherwise could be; interventions like safety cases and monitoring have significant costs. Pacing will be well worth this cost; no amount of American competitive pressure should justify recklessness, or let capabilities get ahead of alignment and monitoring. Where we will need the help of our government is for international coordination. But first we should do what we can ourselves.
4,724
2,182
21,405
4,689,152
Many things that were infeasible in February 2020 were inevitable by March 2020
i think people dismiss the feasibility of international coordination / slowdown on ai far too easily. it doesn’t seem easy or the default outcome, but it seems tractable, in everyone’s interest, and far less “head in the clouds” than the other good timeline alternatives.
7
1,440
Geoff Penington retweeted
Replying to @monke_io
i am never, ever sacrificing my own life to bring another species into existence and i think anyone attempting to do this should get arrested for crimes against humanity. yes, i am speciesist
6
3
238
8,054
Geoff Penington retweeted
the fastest way to lose the frontier will be when, due to the reckless commercial pace of building superintelligent minds, Americans impose a total butlerian jihad. CoreWeave includes this in their risk reporting. you will soon come to see all of this is a moderate solution
Hard not to see “pace the frontier” as “lose the frontier”
159
115
2,190
230,703
Geoff Penington retweeted
WIRED reported today that I began to cry while talking about AI and mathematics. But the article didn’t explain what moved me. The truth is, I’m not entirely sure myself. I’d like to try to explain. 🧵 wired.com/story/mathematicia…
41
242
1,208
248,316
This is a really great statement. I’m also really glad to see @sama immediately endorse it. OpenAI and Anthropic are far more similar organisations than either side sometimes wants to admit, and the future of the planet depends on us being able to work together to get this right
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
4
57
3,510
Geoff Penington retweeted
Dario is right
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
6,350
6,629
58,215
12,567,448
Geoff Penington retweeted
Replying to @sriramk
there should be one that's outside of Berkeley lol
28
10
547
15,945
Geoff Penington retweeted
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
5,173
7,206
67,653
17,012,381
Geoff Penington retweeted
I know it's hard, but when reading try to be generous towards the author. E.g. Fields medallists are not reactionary ego-driven luddites, but think a lot about mathematics, and are sincerely trying to do what they think is best for the subject they have devoted their lives to.
89
84
971
74,076
The ideal scenario here is that Cordoba and Martinez-Zoroa choose to write a proof of Navier-Stokes blow up (either based on the OpenAI construction or otherwise) and then receive the Millennium prize for all of their efforts on the problem
This seems obviously correct and good. Credit is inherently pretty stupid but to the extent it has to exist it should go to a) important human ideas (eg Cordoba-Martinez-Zoroa for Navier-Stokes) and b) human digestion and understanding
3
49
12,073