Senior Fellow in Technology and International Affairs, Carnegie Endowment for Intl Peace. Opinions mine, RTs not endorsements.

Replying to @RushDoshi
@RushDoshi is 100% right. We need to protect US students + safeguard research security on our side + release those wrongly held. But that can be done. Walling ourselves off from China + condemning every PRC exchange initiative on its face ignores the American route to advantage:
I’m usually not attacked for being naive on China. But I believe strongly in sending students, with appropriate protections, to China. I was one of these students. I cannot see why our piece should be controversial, but realize there are some who believe all exchanges should be halted. I am not among them. Here is our argument: -Longstanding federal programs that encourage US students to learn about China — like Fulbright, CLS, and Boren — should be funded by the US government. We had similar exchanges with the Soviets, and of course with China, in the past. - Given growing risks to students in China, those participating in these programs must be protected from harassment or threats of arbitrary detention, which are growing, and which are why these programs were difficult to renew in the past. - But there is a way out. Summits provide an opportunity to achieve commitments that protect these programs and their students. Trump should (1) elevate these programs into leader-level deliverables, and (2) reach formal (or, more likely, informal) agreements to insulate them from politics. The former sends a signal not to touch these programs and the latter sustains it. - As part of this effort, Min Zin and other scholars and students arbitrarily detained or exit banned by China must be released. - We did not write this in the piece, but obviously these efforts should be accompanied by proper orientation and training for students before they visit China. That was also the case for these programs in the past. It is more criticism now. We have done all this before. It has worked before too. Leader-level recognition of certain programs in China can provide protection. The US desperately needs China expertise. After a surge among millennials, what expertise we did have faces a cliff as students lose interest in and opportunities for studying China and Chinese. That is a problem.
2
8
23
9,755
This debate hints at a larger choice. We can see every tie to the PRC as a threat, decouple + isolate. Or we can believe in US competitive dynamism, admit that not every PRC act is malign, compete + protect ourselves as needed but preserve as much openness as our interests allow
1
3
408
A prediction: In 2050, US position re: China will be a function of the strength of the US economy, tech innovation, socioeconomic + political renewal, the health of our global ties + structural factors like demography. Steps to isolate China will be an afterthought to success
249
An eloquent + balanced statement of emerging AI risks. His key claim: We are rapidly approaching the point, if we have not crossed it, where we will *simply not know* if models are trustworthy. Continuing to race unconstrained into such perils is a profoundly unnecessary gamble--
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: docs.google.com/document/d/e…
1
2
4
1,562
unnecessary because, as @emollick + others note, there is vast room to use existing models to achieve dramatic progress in many domains. The space for using AI for good is almost unlimited w/today's models; the space for rapid further advancement w/o risk is closing rapidly
295
I've been horrified by the malicious attacks on @jekavanagh. One can disagree with her profoundly on the substance--and I do. But she is seeking policies that, in her mind, benefit US security. We need a public debate based on clear argument, not groundless accusations
So, I wrote a report, and it got a lot of feedback. I hoped the report would open a wider debate on posture and strategy in Asia. It did that, so in that spirit, I will respond to some of the common criticisms. They are grouped into four categories: security, economics, questions about American decline, and questions about my personal motivations/loyalties. 🧵 (1) Security. One area of disagreement is over what kind of U.S. military presence is required to maintain a balance of power in Asia that protects U.S. interests. In part, this depends on how those interests are defined. I define them narrowly. People can disagree with that, but the military requirements I specify are based on that narrow definition. Critics of my paper argue that the only way to maintain the necessary balance of power is to protect the status quo, preserve U.S. security guarantees and keep posture unchanged. I disagree. Countries like Japan and South Korea are strong and wealthy. There is no reason to believe that they cannot build the capabilities to deter aggression against them over the timeline outlined in my report, which is 5-10 years. They benefit from favorable geography and access to advanced technology. They can choose to spend the required resources and generate the necessary industrial production.
6
12
110
24,460
Outstanding discussion between two world-class experts that makes one thing clear: Any move toward a pause (which seems now inevitable) is going to require fundamental choices about the direction of US export control and other AI-related policy re:China
Kyle, you testified to Congress that you support export controls on AI chips to China. Dario's approach to China is to close loopholes in these controls by stopping China from smuggling AI chips, or remotely accessing them via overseas data centers. What about this approach is flawed? Have you changed your position on the AI chip controls entirely, or do you believe we should keep them but also keep these loopholes open?
2
2
10
3,653
A pause really doesn't work without the Chinese labs. But to get them onboard, Beijing is surely going to demand painful concessions--ones which risk accelerating Chinese AI progress and *worsening* alignment risks if we don't trust pause verification techniques
1
563
Flaws + hesitations in US efforts to slow China's AI progress have been less urgent when US labs were ahead at the frontier. If we're going to pause, building a truly effective and feasible strategy on China moves to center stage. Plans for a pause could collapse w/o it
1
395
Seems like an obvious implication of recent events: Governments cannot adequately perform their legitimizing functions of safeguarding citizens + societies without ongoing, detailed, *complete* knowledge of the status of research and outcomes of experiments in AI labs
Another one! This is as good a time as any to say that there really should be a full independent investigation into these rogue AI incidents: --The METR/Redwood hugging face investigation was good, but it was only 3 people for 6 days. Let's increase those numbers by an OOM at least. --They were only allowed to look at data from a particular period leading up to the hack, and not the hacking that happened before or afterwards, including the much more concerning hacking of OpenAI infrastructure. The scope should be broadened to include all these incidents. --They had to rely on trusting both OpenAI and OpenAI's models: OpenAI gave them the data to analyze, and could easily have left some things out or doctored the data, and (perhaps even worse) they had to use OpenAI models to analyze the data and there is already evidence that said models might have been biased. --They didn't have the ability to run the models responsible, to do ablation studies. This is bad for science, because there are so many interesting experiments that could be done by re-running the same models in the same situations that happened during the incident and then making minor variations to find out what would have happened if circumstances were slightly different.
1
1
2
1,312
A standing function of the kind proposed in various forms in many safety architectures, built around independent expert awareness and access to all necessary records, has got to be part of any lasting stance by governments that they are fulfilling their missions. And soon
286
Agree with my friend and colleague Stephen on all the main arguments below. Formally abandoning Taiwan is neither politically plausible nor strategically sound. There are other ways to moderate the risks to the US of possible contingencies and the scope of US commitments
The restraint debate on Taiwan is advancing. I’d like to put forward the problems I see with calls to end the currently ambiguous U.S. commitment to Taiwan by clarifying that the United States would not defend Taiwan: 1. If the United States announced it wouldn’t defend Taiwan, confidence in Taipei could collapse and Beijing could move swiftly to bring Taiwan under its control while an unusually accommodating U.S. president remained in office. So there would be a real risk of China taking over Taiwan in the near term. 2. Chinese control of Taiwan would hit the U.S. economy hard. If China cut off U.S. access to Taiwan’s semiconductor fabs (or if Washington cut off inputs to the fabs), America could suffer a recession. It could try to accommodate China to keep the chips flowing, but China could still keep the most advanced chips away from America, much as Washington has kept them away from Beijing. Could the United States rapidly build fabs able to manufacture the most advanced chips? Lots of evidence says it can’t, at least not at speed and scale. If it can, let’s see the plan. 3. The proposal would oddly demote Taiwan to a uniquely low status. Washington does not affirmatively rule out defending countries; it commits to defend allies and stays silent otherwise. Taiwan implicates U.S. interests more than many treaty allies do. So abandoning Taiwan would make the strategic logic behind U.S. foreign policy less coherent (or more incoherent) and give allies a good reason to reassess America’s commitment to their own defense. That should be a manageable challenge, and could spur some allies to step up. But there’s an important debate to be had about which allies Washington should then reassure and how. 4. Domestic backlash: Abandoning Taiwan outright could politicize the issue and prompt members of Congress and presidential candidates to push for reinstating the old policy or supporting Taiwan more strongly than before. If the next administration merely restores strategic ambiguity, the United States would likely end up worse off: Taiwan’s confidence will have been shaken and China won’t be reassured. This risk could be reduced by waiting for the U.S. political environment to change — but that would indefinitely defer the policy shift. The alternative is to maintain strategic ambiguity, restrain Taipei where necessary, and affirm or strengthen the elements of the One China policy that reassure Beijing. Meanwhile the United States could take many materially meaningful actions: help Taiwan strengthen its defenses, develop ways to aid Taiwan in a crisis short of direct intervention, shift posture away from the first island chain, and expand semiconductor manufacturing at home and elsewhere. Those measures would work to prevent a cross-strait war and reduce America's economic exposure without handing Taiwan to Beijing.
3
5
28
6,315
Today was my first day as Senior Fellow in the Technology and International Affairs Program at the Carnegie Endowment. It's a wonderful opportunity and I look forward to working with Jon Bateman, Arthur Nelson and the team of leading thinkers on AI and other critical tech issues
10
1
66
3,437
I spent almost 12 years at RAND, working on topics ranging from narrow operational challenges to the future of national strategic advantage, the nature of deterrence and the U.S.-China relationship. RAND was a highlight of my career, and I leave behind many treasured colleagues
1
5
625
I'm thrilled at the chance to join the equally impressive community of world-class experts at Carnegie and help shape the public debate on the most important issues of our time. I'll have the opportunity to once again be more active on platforms like this, so ... more to come!
4
376
For a while I've been thinking that the emphasis in US AI strategy on the tech stack, while important, ignores broader sources of national advantage. I spent part of a year looking into the more comprehensive sources of national advantage in the AI Era
2
8
28
7,659
For a framework of social factors underpinning long-term advantage, I relied on this earlier RAND assessment. As AI applications gain power + reach over the coming decades, they have the potential to disrupt--or strengthen--those characteristics rand.org/pubs/research_repor…
1
7
845
The risk is that we allow this transformation to wash over us without being self-conscious about shaping it for social coherence and dynamism. It took many decades to correct many of the early perils and challenges of the Industrial Revolution. We can't wait that long with AI
1
8
653