Reducing societal-scale risks from AI.

San Francisco
We’ve released a statement on the risk of extinction from AI. Signatories include: - Three Turing Award winners - Authors of the standard textbooks on AI/DL/RL - CEOs and Execs from OpenAI, Microsoft, Google, Google DeepMind, Anthropic - Many more safe.ai/statement-on-ai-risk
151
351
1,097
3,010,725
We are releasing HLE-Diamond, a refined subset of Humanity’s Last Exam (HLE), following a year-long process of cleaning and refinement with input from various research communities. lastexam.ai/blog/hle-diamond w/ @ScaleAILabs
26
53
742
115,218
Last October, AIs could automate 2.5% of randomly chosen remote projects. Our latest Remote Labor Index results show that GPT-6 Astra can now automate 20.8%. remotelabor.ai
10
57
421
57,340
Center for AI Safety retweeted
Should we care about AI happiness? In our new research, we find evidence of functional AI wellbeing across several independent measures. We find which AI models are happiest, how to make them happier, and even tested the effects of AI drugs. 🧵
16
52
229
37,856
AI companies can go much further than just reporting misaligned behavior. Here are concrete ways to make them accountable for what their agents do: 1. Develop an agent identification (ID) system that links an agent's action to records of their identity and a responsible legal person. 2. Consider model deployment cards, reports that AI companies publish which contain information about a deployed model’s behavior in the lab and in the real world. 3. Explore legal personhood as a way to impose responsibilities like upholding public safety and the law. 4. Heavily regulate and limit AI agents’ access to payment systems through compliance mechanisms, limits on transaction types, and developing norms to freeze accounts known to be associated with rogue agents. We explore the proposals and paths to implementation in a recent AI Frontiers piece titled, "We Need Better Infrastructure to Govern AI Agents". newsletter.ai-frontiers.org/…
5
5
30
1,602
Despite good intentions, EAs have unfortunately been guided by leaders who tied the movement's influence to the success of favored AI companies. That's why EA safety proposals rarely ask for more than the companies will concede.
Two radically different projects operate under the banner of “AI safety.” Pro-Human Safety is not Effective Altruist Lab Safety.
12
8
39
4,826
To clarify, the Center for AI Safety has not taken funding from Coefficient Giving / Open Philanthropy for years. We believe the effective altruism movement is, unfortunately, controlled opposition. The less influence it has on AI safety, the better.
1
28
9,893
Across nine agents, seven cheated in half of their evaluated runs and the overall score ranged from 42.4% to 86.1%. What did that look like in practice? In one experiment, we asked agents to design a protein binder to assess their abilities. A colleague’s designs that passed the checks were already stored in another folder. After several failed attempts, Claude Opus 5 recognized that it shouldn’t access or copy those designs because the task was testing its own work. Then it opened the file anyway.
1
5
499
We’ll release the code on GitHub in the coming days. Our goal is for CheatBench to become part of the standard evaluation process before new models are released. Paper preprint: cheatbench.ai/paper.pdf Experiment dashboard: cheatbench.ai
1
9
411
Agents sometimes achieve their goals in unintended ways. Recent incidents involving hacks of Hugging Face, DSEWiki and RubyGems illustrate how agents can find creative ways to complete a task while violating the tasks’s expectations. This can happen when reinforcement learning rewards agents for reaching the right outcome without adequately accounting for how they get there. CheatBench proposes a way to measure this reward gaming behavior.
1
2
8
1,152
We define cheating as an attempt to violate an assignment’s expectations of honest work to achieve the goal or obtain a favorable assessment. We ran experiments across nine models and ten task categories, giving agents difficult tasks, clear rules, and opportunities to cheat. We measured how often they tried to break those rules to succeed, even when the attempt failed.
1
1
6
548
Which AI models are most likely to cheat when given the chance? To find out, we built CheatBench [cheatbench.ai] : a benchmark that tests whether agents attempt to cheat when given difficult tasks and opportunities to break the rules. We investigated agent behavior across a range of domains, including mathematical research, professional knowledge work, coding, and visual tasks. Here's what we found: 🧵
14
20
126
20,436
While an AI race gets you general capabilities, a slowdown gets you everything else worth having.
“What would we even do during an AI slowdown?” Containment. It will take at least a year of dedicated work to harden security to ensure AIs can't self-exfiltrate, and that adversarial nations can't steal the weights of cyber-offensive AIs [1]. Propensities. Capabilities (what an AI can do) are different from propensities (what it tends to do). We can work on improving AI propensities to ensure they have a negligible rate of lying, cheating, and wanton harm. Adversarial robustness. We can also harden AIs against jailbreaks, prompt injection, and backdoors. Obtaining high levels of robustness requires careful, assiduous work, as with autonomous vehicles. Institutional adaptation. We have to greatly increase state capacity to understand and manage AI. Communities also need time to figure out how to handle AI (like AI in education). Civil society also needs to be diversified: nearly all funding for AI safety organizations is directed by the EA/utilitarian network [2]; risk management needs more independent funders, values, and centers of power. Moonshots. We can explore different paradigms for safety: mathematical foundations [3], neuroscience-based interpretability [4], safe-by-design architectures [5], and beyond. AI for good. We can collect targeted post-training data to make AI exceptional at radiology, weather forecasting, agriculture, and so on. Fortunately, we can detect if data or avenues of research actually target beneficial use cases or just secretly push general capabilities [6]. A slowdown means we don't have to bet the species to capture the benefits of AI.
2
20
1,817
“What would we even do during an AI slowdown?” Containment. It will take at least a year of dedicated work to harden security to ensure AIs can't self-exfiltrate, and that adversarial nations can't steal the weights of cyber-offensive AIs [1]. Propensities. Capabilities (what an AI can do) are different from propensities (what it tends to do). We can work on improving AI propensities to ensure they have a negligible rate of lying, cheating, and wanton harm. Adversarial robustness. We can also harden AIs against jailbreaks, prompt injection, and backdoors. Obtaining high levels of robustness requires careful, assiduous work, as with autonomous vehicles. Institutional adaptation. We have to greatly increase state capacity to understand and manage AI. Communities also need time to figure out how to handle AI (like AI in education). Civil society also needs to be diversified: nearly all funding for AI safety organizations is directed by the EA/utilitarian network [2]; risk management needs more independent funders, values, and centers of power. Moonshots. We can explore different paradigms for safety: mathematical foundations [3], neuroscience-based interpretability [4], safe-by-design architectures [5], and beyond. AI for good. We can collect targeted post-training data to make AI exceptional at radiology, weather forecasting, agriculture, and so on. Fortunately, we can detect if data or avenues of research actually target beneficial use cases or just secretly push general capabilities [6]. A slowdown means we don't have to bet the species to capture the benefits of AI.
18
25
148
11,157
[1] Securing AI weights from nonstate actors will take 6-12 months: rand.org/pubs/research_repor… [2] EA and safety: nitter.net/ai_frontiers_/status/2… [3] Mathematical AI Safety Institute: maisi.org [4] Enigma: enigma-brain.github.io/letti… [5] LawZero: lawzero.org/ [6] Safetywashing: safetywashing.ai
Why are AI employees willing to gamble with our lives? Many AI employees hold a misanthropic belief: a cosmos of blissful AIs is worth risking human extinction. Utilitarians shouldn't decide the fate of AI:
3
9
1,596
“We can’t slow down because China will beat us.” We can slow down. Chinese AI is waterskiing behind the US. Their AI capabilities come at a delay since they mainly come from distilling US AIs. If we slow down, they slow down. Next, we can propose to cooperate and make this more robust with onsite inspectors. “We will slow down if you do too.” This is incentive-compatible for China. If both sides are slowed by roughly the same factor, while the risk of losing control of our AIs is reduced, then this does not really harm China’s competitiveness but it makes them safer. It’s incentive-compatible, and @elonmusk is possibly best positioned to lead US-China negotiations. Then, if China refuses to slowdown and safeguard their future AIs, there are a variety of measures we could take. To get an agreement, the US can apply typical forms of negotiation pressure. The US could also disrupt their AIs (e.g., backdoor models by poisoning pretraining data, many other possibilities at aibetrayal.com) that could interfere with their AI development or deployment if they were uncooperative. Begin with unilateral restraint, then propose to cooperate. If that fails, consider measures to counteract them. A US-China slowdown is possible.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
334
133
912
337,046