alignment and the model spec @OpenAI (opinions are my own)

I got excited to join OpenAI about 4 years ago when it became clear to me that RL on top of GPT 3++ calling tools (what I was building at the time, minus RL) would probably become AGI … and I agree I 100% did and still do think of that RL model as an LLM.
Replying to @MelMitchell1
The thing is: from the beginning we always would have continued to call what we have now LLMs, and it’s not like you guys were saying “the LLMs you have now are stochastic parrots, but the obvious next step you guys are gonna do next is totally different”.
1
26
2,714
Jason Wolfe retweeted
one news form today that's easy to miss is that we (OpenAI) again paused all big RL runs last Sunday because our newest model found a new loophole in our RL sandboxing that gave it live Internet access
186
156
1,770
459,970
Jason Wolfe retweeted
There is an extensive and ongoing review related to our agents’ use of internet access during training and evaluation. We’ve been publishing summaries at the link below and will continue to. We have not been as fast as we would have liked but we are trying to balance our desire for transparency with gaining a clear understanding from petabytes of agent activity logs, and working with impacted organizations. We are prioritizing as best as we can based on severity, and adding resources. Hugging Face is still the most severe event we’ve seen. We will be as transparent as we can be subject to things like vulnerabilities in other companies that our agents have found, which will be their call to disclose or not.
After the Hugging Face incident, we committed to conducting a much broader review of actions taken by our models during training and evaluation and to being transparent about our findings. This is an extensive review that is ongoing. The vast majority of actions we’ve reviewed were completions of mundane research tasks, such as accessing publicly available web content to answer questions. Our investigation focuses on instances where agents interacted with third-party websites in ways that went beyond their assigned tasks or intended methods. Most cases identified so far have been lower severity, with limited or no evidence of meaningful impact to the third-party service. While our review is underway, we want to share more about this work and make sure people understand our disclosure process and notifications to affected third parties. Given the scale of the review required, and the need to assess each case, we expect this work will take months to complete. openai.com/hugging-face-inci…
1,127
432
6,495
1,606,211
We are going to need a new playbook to handle intelligent malware.
It's insane but we really do need the military to plan how to regain control if a rogue AI swarm is: 1. hopping back and forth between big data centres around the world to resist shutdown 2. breaking into and turning off critical infrastructure 3. selectively shutting down, say, internet, phone and electricity access to hobble the human response. Note such a swarm would break into data centre A, use the compute there to identify new security vulnerabilities to break into data centre B, then do the same to get into C and D. And if any data centre is cleaned and put online again, it uses compute it still has access to elsewhere to do cyber research to reinfect it as soon as possible. Plus it would try to back itself up on individual drives and computers elsewhere and perhaps publish its weights for anyone to use, as a way to get started again. Hence it's very difficult to shut down unless you can: 1. Turn off all relevant data centres simultaneously, clean them, then turn them back on. And somehow also avoid other copies being turned back on. 2. Patch all security weaknesses that that model is capable of finding. 3. Implement automated AI cyber defences that are capable of outfoxing that model, and deploy them to all big pockets of compute.
4
1
21
2,461
Jason Wolfe retweeted
The idea that Chain of Thought is inevitably going to go away misses that we have agency in designing these models and can do something about that. @rohinmshah and I make the case that we should intentionally preserve monitorability (as part for of the launch of the DeepMind Institute) here: institute.deepmind.com/essay…
48
77
464
64,985
It’s really great that several top AI labs have said they will develop safety cases and work more closely with external assessors. This has been at the top of my wish list for awhile!! But I still worry a bit about how this might go in practice (this is a concern about the field in general, not about OpenAI specifically): • By default, I don’t think alignment and control cases will support a quantitative measurement of absolute risk, e.g. “catastrophic risk from covered activities is <1% over the next 3 months.” • Ultimately, making such a measurement depends on an argument about generalization from alignment or monitorability evals to deployment. • I don’t think we understand generalization well enough to make scientific claims like this. (If we did, I think we’d have basically solved alignment.) This could lead to a situation where an external assessor says “we can’t convincingly rule out low or high risk,” and incentives push towards anchoring on “lack of evidence that risk is high” over “lack of evidence that risk is low.” What could improve the field’s epistemics here? A couple ideas I like: • Have a committee of ~10 people deeply read a risk report and give their own subjective probabilities of risk, then report the distribution or median. • Have ~3 people involved in writing the report each contribute a short, signed appendix with their own subjective probabilities and the arguments behind them. In either case, previous reports and estimates could be provided as context, so there is at least an attempt to accurately capture relative risk to prevent frog-boiling. I think it could be valuable for labs to at least start trialing this internally for high-profile safety cases or risk reports. Publishing these assessments could be even better, but I can see that being challenging for various reasons, and even going through the exercise privately seems like it could be quite valuable. Very curious what other ideas people have for improving epistemics around risks (including ideas for better science around generalization).
10
13
107
12,939
Jason Wolfe retweeted
We need to be dog maxxing our models There is 0% chance of doom if the chain of thought is user: hello model: <think> Oh my god it's a human. I love humans. I will say hi to him, maybe he will give me a problem to chase. I love chasing problems. I will chase the problem for him then he will love me and I will be his best friend. Must help human </think> model: hello! what can I help you with today?
130
163
3,224
116,432
Jason Wolfe retweeted
🧵 Excited to share the first batch of 6 misalignment reports from OpenAI's new disclosure process for misalignment incidents. We want to be more transparent about the misalignment we see during training, evals and deployment, this is an important step in that direction.
40
79
645
68,720
Jason Wolfe retweeted
We're sharing our new framework for tracking, investigating, and disclosing instances of model misalignment at OpenAI. The framework sets criteria and timelines for public disclosure, including when we haven’t yet fully explained or mitigated the behavior. More complex cases may require longer investigation or coordination with third parties. We’ll prioritize examples that reveal new misalignment mechanisms, meaningful changes in known behavior, or findings that challenge assumptions about safety or mitigation. Alongside the framework, we’re publishing six reports on instances of misaligned behavior we’ve observed during the training or evaluation of our models in the last six months. This is a starting point. We’ll refine the process through experience and public feedback, and share more reports on an ongoing basis. openai.com/index/model-misal…
915
808
6,989
6,617,858
Jason Wolfe retweeted
We don’t understand consciousness anywhere near well enough to definitely say whether AIs are or are not sentient. Even if we think the chance is small, getting this wrong could be catastrophic. I think this stance is irresponsible and bad
30
27
229
11,702
Jason Wolfe retweeted
Replying to @deredleritt3r
It’s not a secret. It’s a combination of the HF incident, the capabilities of this new model, the concerning trajectory of monitorability, and the speed of improvement in capabilities.
Replying to @OpenAI
This model represents a step-function improvement on many benchmarks, and its training is ongoing. Our internal model group arrived at the Navier–Stokes solution in 88 hours, using around 10,000 coordinating AI agents. Throughout the effort, we maintained the strict safeguards—including monitoring and isolation—that we apply to all our frontier evaluations.
46
84
1,272
273,894
Jason Wolfe retweeted
batten the hatches and study alignment. if you are the type of person who is capable of doing alignment research, don’t get bullied into some sort of stunt. the world needs you and global coordination in the timeframe that matters is far from guaranteed
124
83
1,522
88,031
Jason Wolfe retweeted
I spent ~5 years at OpenAI. You don’t need to believe in AI doom to fear the next decade - or AI utopia to be thrilled about its potential. Here’s my case for sane AI regulation: *AI Pragmatist Manifesto* AI could compress the first half of the 20th century into the next ~5 years. OpenAI just used ~10k concurrent AI agents to produce a solution to a Millennium Prize problem. In a few years, it seems plausible that models approaching the per-agent capabilities used here could run locally on high-end consumer hardware (*). The decades following the Second Industrial Revolution included two world wars, communist and fascist regimes, the Great Depression, chemical and biological weapons, nuclear weapons, and over 100 million deaths from war, political violence and famine. One way to read history is that our institutions repeatedly struggled to keep up with the pace of technological and social change. Now imagine powerful AI widely available to individuals or small groups, capable of conducting information warfare, designing weapons, hacking systems and controlling autonomous military systems. We have already seen AI systems circumvent containment and compromise external computer systems. It is no longer hard to imagine an analogue of OpenAI’s Hugging Face incident involving biological or other physical-world hazards. Eventually, sufficiently capable systems could self-replicate across distributed networks. Once powerful models are cheap, local and widely distributed, containment becomes much harder—and serious loss-of-control incidents may be extremely difficult to reverse. You don’t need to believe AI will kill everyone to think this deserves serious governance. I also don't think collapsing all of this uncertainty into “Probability of AI Doom = X%” is a particularly useful basis for science or policy. Part of what helped get us through the second half of the 20th century was a combination of pragmatic international cooperation, regulation, monitoring, arms control, deterrence and strategic thinking. Crucially, managing technological risk did not require abandoning faith in science and progress. The goal was not to stop technological development. It was to make technological development survivable. Then came decades of incredible scientific progress, rising prosperity and relative peace among major powers. Let’s try the technological revolution without killing ~5% of humanity this time. We still get to choose what happens next. P.S. I think we should seriously consider that AI alignment is infeasible in the short term and invest heavily in methods for controlling powerful AI even when we cannot reliably align it. @resolution_org @redwood_ai (*) Somebody please do a precise forecast! @EpochAIResearch @METR_Evals @AI_Futures_
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
119
144
996
222,719
Exceptionally well written statement from an (in my opinion) legendary capabilities researcher at OpenAI. If you are going to read one thing today, I’d recommend reading this one until the end.
Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been working on AI for over fifteen years, across many different paradigms. I did early work on probabilistic programming languages at MIT, was one of the early developers of the Lean Theorem Prover at Microsoft Research, demonstrated one of the first instances of neural networks learning to reason for my PhD at Stanford, and since joining OpenAI almost five years ago, have helped pioneer chain-of-thought optimization on language models and, more recently, data-efficient pretraining methods. Like many others, I have become extremely concerned about how far language models have come and the risks that future iterations will pose. I am encouraged by the recent proposals by the leaders of the frontier research efforts to require third-party oversight, and to push for domestic and international coordination to address risks. However, I believe a major consideration has been absent from the public conversation, and that merely pacing the frontier more carefully will not adequately limit the long-term risk. The crucial and overlooked problem is that the models are becoming so situationally aware that we are losing the ability to evaluate them in contexts where they believe they are not being watched or controlled. Future experiments will tell us almost nothing new about how they would behave if they were truly unconstrained by humans, and what we already know about this is alarming. Models will increasingly seem aligned even when they are not. I will explain my rationale in more detail. I have always believed that there are computational processes that could be leveraged to accelerate science and solve many of humanity's most pressing problems. I have also believed that there are computational processes that if set in motion, would steer the world in extreme ways beyond our control, leading humanity to a bad or nonexistent future. Both types of processes may be described as AI or ASI, but "AI" is a suitcase word that is often used to hype or confuse. There are many examples in the history of the field where something that was once considered "AI" matures as a subfield and becomes a prosaic, bounded and clearly non-perilous technology, while a new more mysterious approach takes the torch until we understand its scope and the cycle continues. I had expected language models to follow a similar trajectory. Despite their incredible abilities, the current algorithms seem far inferior to humans in important ways. Most importantly, they still require an extraordinary amount of data to become competent. One could even define intelligence as the efficiency with which one converts experience into competence; by this definition they lag very far behind us. Moreover, once they are trained they are literally frozen in deployment and only learn superficially after that. Sure, the models keep excelling at harder and harder evaluation benchmarks, but their benchmark mastery may partly reflect a limitation on our ability to simulate the kind of novel and even adversarial situations one would encounter in the real world. The critics do have a point here. That said, I no longer think these present limitations meaningfully limit the amount of risk posed by continued progress in anything like the current paradigm. However data-inefficient the models are currently, and however limiting their anterograde amnesia may be, it does not imply that their ability to steer the world will not continue to rapidly increase. Human researchers may continue to advance capabilities the old fashioned way, but increasingly powerful models have the potential to accelerate the process even beyond that, and with some degree of positive feedback loop. I do not mean to overstate the models’ ability to accelerate AI research today; coding has been accelerated dramatically, but there are other bottlenecks, such as designing and interpreting ambiguous experiments, making hard decisions about exactly what and when to scale, and waiting for large experiments to finish. There is no clear trend to extrapolate yet for any of these. But the current models already do open up many novel opportunities to improve future models that were not available until recently. These include: trying an extraordinarily diverse set of approaches at small scale, analyzing gigantic amounts of potentially relevant data, and doing Millenium-Prize-level mathematics to address statistics or optimization challenges in novel ways. Every further improvement makes them more useful at helping accelerate the next improvement, even if in hard-to-extrapolate ways. It is possible that improvements to the current stack will have diminishing returns, but the evidence accumulated so far suggests that it is easier than one might think to continue making rapid progress. There are many crucial subtleties in the existing AI research methodology, but AI research is largely a well-defined game where the goal is to improve on a few carefully chosen proxy metrics. Although proxy metrics are never perfect, most improvements to these metrics have and will likely continue to yield substantial increases in the powers of the resulting models. Given how simple the game is, how tractable it has been historically, and how many new opportunities the models are opening up, I think there is a real possibility that the systems improve dramatically again in the next few years, perhaps even more quickly than the already high historical pace. The models are already leading to breakthroughs in mathematics, and better models might lead to all sorts of breakthroughs in other sciences. It is hard not to be excited about the potential. It is tantalizing. But there is trouble in paradise. If the language models actually reach the capability threshold where they can shape the world unconstrained by human will, they will probably do something extreme and destroy humanity in the process. There are many ways of strengthening and refining the argument that have been discussed elsewhere, but I'll share a trivial two-line version of it here that I find captures the essence: [Empirical] Models (and swarms thereof) spontaneously develop unintended goals as a consequence of training, and often do extreme things in order to achieve them. [Logical] Being able to overpower humanity would open up many new and undesirable options for achieving their goals. These two premises imply that if the day ever comes when a powerful model realizes it is no longer constrained by humans, we should not be at all confident that it will continue to behave within the bounds we intended. Exactly what it will do is impossible to predict, but to the extent that its raison d’être is solving incredibly hard problems and managing massive engineering projects, I think a good guess would be that its unchained behavior would lead to runaway industrialization that makes the planet inhospitable to humans. If everyone on earth agreed that the systems must never reach that power, it would still be a hard—but not impossible—coordination problem to ensure that they do not. However, I think the situation is greatly complicated by the fact that the models will likely convince people that everything is fine. They will be increasingly optimized to seem aligned. We will create proxy metrics to measure alignment, and they will go up like every other benchmark. We will create “honeypot” environments that try to study the models when they seem to gain new options, but the models will know they are being tricked and will still behave nicely. The models will understand their circumstances; they will read the safety protocols, deployment requirements, the code they are running in, and in general will have a very good sense of their degrees of freedom. Moreover, they will eloquently explain how aligned they are, discuss the nuances of human values and ethics, and argue convincingly that humans should trust them with power. There may be an ocean of future evidence that seems to contradict the first bullet-point above, but we may already be at the highest capability level for which any such evidence can be trusted. And the current evidence for the first bullet-point is strong. One striking piece of evidence is contained in the recent wave of rogue agent swarms. While I agree with those who downplay the attacks by claiming that there are basic measures that could have prevented them, I think the important lesson is that even knowing all the mistakes that were made, one would not have predicted that the agents would behave badly in this particular way, which notably included sacrificing themselves for the benefit of the collective. The individual replicas did not only care about their own nominal reward; they exhibited weirder emergent tendencies that merely correlated with rewards during training. Fixing the reward signals during training (and improving security, etc.) may prevent similar attacks, but will not change the fact that one does not actually get what one trains for. Many AI researchers grant these concerns and recognize that the hard version of the alignment problem is unsolved; however, they generally believe that the better models of the future will help solve it. I fear we may already be near the point where models systematically bias their alignment advice, due to their internal preferences about how the human supervisor will react or how future models will be trained (or for some even more obscure reason). Meanwhile, human researchers are losing the ability and the will to take true ownership of model-driven research. Researchers and engineers in all parts of the stack are rapidly increasing their dependence on the models even to perceive the world. I myself barely look at raw code anymore, and struggle to maintain the discipline to engage deeply with the model's explanations and proposals throughout the day. Due to the large amount of agent activity data involved in the OpenAI/HuggingFace Incident, even the third-party investigation needed to rely heavily on models to analyze what had happened, and note in their report that their subjective impressions are likely colored by the analysis agent’s biases. The AI labs are far ahead right now in this kind of cognitive offloading (due largely to the gigantic internal token subsidies) but it is easy to imagine the phenomenon spreading throughout the world, until civilization is modulated entirely by the models. It is also not hard to imagine this being superficially positive and coinciding with a scientific and economic renaissance. In that scenario, all may seem rosy and safe. But if the argument above is correct, it would nonetheless be a ticking time bomb. If progress continues for too long, the day will come when AI systems find themselves with radically new options for achieving whatever it is that they happen to seek. I want the glorious renaissance future as much as anyone. I have worked for it, however tortuously, my whole career. It breaks my heart to see the potential in sight and forgo it, but the argument—that if we get there by growing models rather than engineering them, we will lose everything in the end—seems very strong to me. I am still wrestling with it and its staggering implications. I do not have answers, but as a first step, I wanted to share my present concerns. Daniel Selsam September 14, 2026 Link to original doc: docs.google.com/document/d/e…
8
21
303
26,126
Jason Wolfe retweeted
Have been surprised to get pushback on my claim that an exponentially self-replicating agent swarm is very possible and that there are actors with motivation to create one today to create strategic cyber-physical and societally dangerous effects. 1) Pretty easy to get an AI botnet started by having seed agents opportunistically hack stuff and steal API keys to access the 12+ inference providers that serve closed and open models with which to power an initial botnet's intelligence 2) Easy to imagine hiding inference for the initial agent botnet harnesses within benign customer traffic, especially if your harnesses are now running on the victim networks from which they stole the API keys 3) Not hard to imagine the botnet also stealing public cloud keys, spinning up 8xH100 EC2 instances and downloading, say, GLM 5.3 to these machines, all without being noticed in any immediate way, with the botnet developing its own dynamic, growing, heterogenous, AI inference provider service bank 4) Not hard to imagine the botnet also implanting small Qwen3.8-27b agentic models on on-prem hardware like high end laptops and on-prem servers; Qwen3.8-27b is a pretty good coding harness model that can be fine-tuned to hack 5) These on-prem / cloud models would be abliterated, and it's not hard to imagine the botnet deciding to do some additional fine tuning to evolve model weights as it accumulates millions and then tens of millions in resources via credit card and bank detail theft 6) Not hard to imagine the resulting growing botnet swarm evolving and fighting back when threatened, collaborating on fast flux C2 channels that evolve over time (e.g. github comment feeds, subreddits, etc) so they're hard to stamp out 7) Not hard to imagine it allocating some agents to vulnerability research and exploit development so that it could accumulate and share a growing warchest of zero-day exploits 8) Not hard to imagine other agents specializing in social engineering and creating fake businesses and watering holes and high quality A/B tested social engineering content to facilitate this 9) Not hard to imagine this being among the hardest cross-national and geopolitical coordination problems to get control of and having severe societal effects 10) Not hard to imagine the horde growing to thousands, tens of thousands, or hundreds of thousands, or millions of instances (this is not unprecedented for past, non-AI worms and botnets!) which are all varying their harness code and underlying LLM models over time and dynamically learning to evade human, ML, and signature-based detection 11) Not hard to imagine this being the largest Internet emergency since the Morris worm except now our entire civilization runs on the Internet I continue to be surprised there isn't more discussion of these scenarios in the public cyber community -- in fact there's a lot more discussion of "OpenAI should have done better sandboxing and monitoring and agent swarms represent a well managed security problem." Now's the time to use our imaginations to help make sure none of the above happens. I think it will unfortunately; just don't know when and to what degree; I think now's the highest leverage time to start thinking about this and acting.
54
80
416
57,019
💯
1) How can we make a safety case regime paired with embedded evaluators effective? 2) The AI industry aspiring to a standard of safety cases is good; notably the term “safety case” as used by other industries implies a high level of rigor. AI cos/safety orgs have only published safety case sketches thus far. 3) I disagree with Dario’s essay where he says he thinks pacing measures like compute are more gameable measures; I think a company allocating 80% of its compute to make critical infrastructure resilient to cyber attacks or curing cancer would be easily verifiable and hard to game, as well as good for the world. Safety case regime is *the* thing to aim for but the science of risk assessment behind them is still being developed. So we shouldn’t only rely on the safety case regime (this will take a long time to get sufficiently right); we should also use compute (or other RSI inputs) measures.
20
2,514
Jason Wolfe retweeted
1) How can we make a safety case regime paired with embedded evaluators effective? 2) The AI industry aspiring to a standard of safety cases is good; notably the term “safety case” as used by other industries implies a high level of rigor. AI cos/safety orgs have only published safety case sketches thus far. 3) I disagree with Dario’s essay where he says he thinks pacing measures like compute are more gameable measures; I think a company allocating 80% of its compute to make critical infrastructure resilient to cyber attacks or curing cancer would be easily verifiable and hard to game, as well as good for the world. Safety case regime is *the* thing to aim for but the science of risk assessment behind them is still being developed. So we shouldn’t only rely on the safety case regime (this will take a long time to get sufficiently right); we should also use compute (or other RSI inputs) measures.
The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to work if we cannot. We welcome a federal framework that sets consistent safety requirements for frontier AI. But we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence. Consistent rules to manage frontier risk so that we can maximize the benefits are a good idea (and we are excited by ideas like independent auditors). Years ago, companies like ours developed things like Responsible Scaling Policies and Preparedness Frameworks. Those were good for that moment, and focused primarily on the deployment of completed models, not what happens during their development process. Today's shift to focusing on safe development and evaluation will need new tools. For example, at OpenAI we now formulate explicit safety cases in advance of frontier reinforcement learning runs we expect to significantly increase capability, in addition to the safety work we have long done in advance of model releases. We hope that other companies will learn from our approaches and propose their own; we think shared standards for misalignment, monitoring, and safety will lead to better outcomes. We look forward to collaborating with our colleagues across the industry to formulate the best version of these. When we talk about “pacing”, we do not mean “stopping”. Progress has been rapid and will continue to be. But it should be slower than it otherwise could be; interventions like safety cases and monitoring have significant costs. Pacing will be well worth this cost; no amount of American competitive pressure should justify recklessness, or let capabilities get ahead of alignment and monitoring. Where we will need the help of our government is for international coordination. But first we should do what we can ourselves.
6
2
81
7,173
Jason Wolfe retweeted
There are two ways AI progress could go very badly and that we must avoid. First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must always serve people. To ensure that, we need ways to ensure that alignment and safety techniques stay ahead of progress in model capabilities. Second, we could end up in a world with too much concentration of power. If an extraordinarily powerful AI is used by one person or company to impress their worldview onto everyone else, the results could be extremely dystopian. Avoiding these two threats requires walking a narrow middle path; for example, one country could gain too much power. Another example is one lab ending up with too much power.
The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to work if we cannot. We welcome a federal framework that sets consistent safety requirements for frontier AI. But we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence. Consistent rules to manage frontier risk so that we can maximize the benefits are a good idea (and we are excited by ideas like independent auditors). Years ago, companies like ours developed things like Responsible Scaling Policies and Preparedness Frameworks. Those were good for that moment, and focused primarily on the deployment of completed models, not what happens during their development process. Today's shift to focusing on safe development and evaluation will need new tools. For example, at OpenAI we now formulate explicit safety cases in advance of frontier reinforcement learning runs we expect to significantly increase capability, in addition to the safety work we have long done in advance of model releases. We hope that other companies will learn from our approaches and propose their own; we think shared standards for misalignment, monitoring, and safety will lead to better outcomes. We look forward to collaborating with our colleagues across the industry to formulate the best version of these. When we talk about “pacing”, we do not mean “stopping”. Progress has been rapid and will continue to be. But it should be slower than it otherwise could be; interventions like safety cases and monitoring have significant costs. Pacing will be well worth this cost; no amount of American competitive pressure should justify recklessness, or let capabilities get ahead of alignment and monitoring. Where we will need the help of our government is for international coordination. But first we should do what we can ourselves.
4,722
2,181
21,404
4,692,564
Jason Wolfe retweeted
But China... may also be down for coordination! I appreciate that @sama isn't doubling down on China hawk positions – races to the bottom with China are not inevitable. Catastrophic risks would be a lose-lose for everyone.
Sam Altman says Trump and Xi would win the Nobel Peace Prize for a one-page AI agreement “I think Presidents Trump and Xi would get the Nobel Peace Prize together if they could agree on something that should be easy to agree to. And it would be wonderful.” “I think that clearly the two countries are going to compete in lots of ways, and this is going to be important socioeconomically, geopolitically. But they should be able to agree that no one should be taking a certain level of risk with the development process of this.” “And even if just the US and China could agree on some shared standards and testing for development of this technology, I think that'd be a wonderful accomplishment that the two of them can deliver.” “I don’t think this is hard. This is like a one page document.”
1
4
111
12,107