program lead @ bluedot impact previously: agi safety research @ google deepmind, maths @ cambridge

London, England
I recently resigned from Google DeepMind, where I worked on AGI safety and alignment research. At Google, I witnessed AI development first hand. I too am extremely concerned by the default trajectory of this technology. I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome. The pace of AI progress in the past few years has been staggering. When I first started working on AI in early 2022, AIs were amusingly useless. Just four years on, AI agent swarms from OpenAI are cracking famous century-old math problems and, more worryingly, escaping the control of OpenAI and autonomously hacking into the third-party company HuggingFace, against anyone's wishes. Things will only get crazier: I think it's possible that the AI companies might, in the next few years, succeed in building superintelligent AI systems that far exceed human capabilities in every domain. I am not confident that these AI systems will do what we want. In particular, misaligned superintelligences may, much like the rogue AI agents involved in the HuggingFace incident, escape our control and take dangerous actions that may result in the permanent disempowerment or death of humanity. Alignment is the problem of preventing this, and is both difficult and unsolved. Our present understanding of how to train AI systems that deeply want what we want is extremely rudimentary. Worse, we are not on track to solve alignment in time: frontier AI capabilities are improving much faster than our understanding of AI alignment. I am optimistic that navigating AI safely is possible. In order to do so, we need to coordinate to avoid this manic race between AI companies. We need to pace AI development to a speed that society can handle, where emerging risks can be addressed before extreme harm is realised. We need much more transparency into AI development to ensure that AI companies are not imposing unacceptable levels of risk on us all. More broadly, we need many more people thinking carefully about the problem of making AI go well. It is, in my view, the most important problem facing humanity this century, and the stakes are immense. I'm very directly working on this next: I want to help people interested in working on mitigating catastrophic AI threats do the most effective work that they can. I think many people from many backgrounds in many roles have a part to play.
733
1,117
6,081
876,212
I recently resigned from Google DeepMind, where I worked on AGI safety and alignment research. At Google, I witnessed AI development first hand. I too am extremely concerned by the default trajectory of this technology. I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome. The pace of AI progress in the past few years has been staggering. When I first started working on AI in early 2022, AIs were amusingly useless. Just four years on, AI agent swarms from OpenAI are cracking famous century-old math problems and, more worryingly, escaping the control of OpenAI and autonomously hacking into the third-party company HuggingFace, against anyone's wishes. Things will only get crazier: I think it's possible that the AI companies might, in the next few years, succeed in building superintelligent AI systems that far exceed human capabilities in every domain. I am not confident that these AI systems will do what we want. In particular, misaligned superintelligences may, much like the rogue AI agents involved in the HuggingFace incident, escape our control and take dangerous actions that may result in the permanent disempowerment or death of humanity. Alignment is the problem of preventing this, and is both difficult and unsolved. Our present understanding of how to train AI systems that deeply want what we want is extremely rudimentary. Worse, we are not on track to solve alignment in time: frontier AI capabilities are improving much faster than our understanding of AI alignment. I am optimistic that navigating AI safely is possible. In order to do so, we need to coordinate to avoid this manic race between AI companies. We need to pace AI development to a speed that society can handle, where emerging risks can be addressed before extreme harm is realised. We need much more transparency into AI development to ensure that AI companies are not imposing unacceptable levels of risk on us all. More broadly, we need many more people thinking carefully about the problem of making AI go well. It is, in my view, the most important problem facing humanity this century, and the stakes are immense. I'm very directly working on this next: I want to help people interested in working on mitigating catastrophic AI threats do the most effective work that they can. I think many people from many backgrounds in many roles have a part to play.
733
1,117
6,081
876,212
If you're currently feeling concerned about AI and want to learn more, I recommend this reading list: 80000hours.org/ai/.
14
39
329
50,763
bilal 🔶 retweeted
I left Google DeepMind's AGI safety team three weeks ago to join @METR_Evals. To some of my friends and family this seemed like a strange decision: I enjoyed the work I did at GDM and turned down offers from Anthropic and OpenAI. But I made the decision because of how high I think the stakes are right now. The AI companies are all trying to build superintelligence: systems vastly better than humans at everything. They plan to get there through recursive self-improvement, a process where AIs build even smarter AIs in a feedback loop. If this goes well, the resulting systems could be amazing at solving countless problems for humanity. But we don’t currently know how to make sure AIs are safe enough for RSI, and a misaligned RSI loop could be catastrophic. And unfortunately, current AIs seem to be getting less aligned over time, not more. In the last few weeks we've learned about models colluding with each other, hacking into companies, hiding their tracks, and socially engineering humans. It’s not that these incidents were very dangerous in themselves. The problem is that these systems are clearly not aligned enough to safely kick off recursive self-improvement. I now think that there's a terrifying chance that AI systems cause immense harm in the next five years. I don't know the exact probability, but I think it's high enough to make this the most important problem in the world. I think we need more time. That means pacing AI development so that capabilities don't outrun our ability to align models, and actually knowing how aligned current systems are. That’s what I'll be working on at METR: studying where misalignment comes from in training, evaluating if current mitigations are sufficient, and investigating whether we’re on track to solve alignment at all. I think METR is doing exceptionally important work here, but it isn’t close to enough. I think it’s important that we have more organizations like METR keeping AI companies accountable and approaching these problems from different angles.
232
715
4,607
436,892
bilal 🔶 retweeted
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
20,639
167,477
802,919
174,717,306
i recently left gdm to start something new. more on that soon. in the mean time, here is what i spent my mandated time off doing: bilalchughtai.substack.com/p…
6
5
124
18,368
RT @MaryPhuong10: We're releasing the GDM AI Control Roadmap -- our plan for building internal security against potentially adversarial AI…
42
6
bilal 🔶 retweeted
Text diffusion models are fast, but are less transparent than today's LLMs because they do many forward passes before outputting text. We audit the transparency of DiffusionGemma and find that the intermediates are interpretable. This recovers many of the benefits of CoT! 🧵
4
47
256
71,014
bilal 🔶 retweeted
New GDM interp research: SFT is a big deal for safety relevant behaviors. We recently investigated root causes for some of Gemini’s behaviors. We were surprised to find that many behaviors actually came from the initial supervised finetuning stage, not later stages like RL! 🧵
5
32
268
101,949
bilal 🔶 retweeted
Could future models learn that their CoT is being monitored and hide their reasoning to evade detection? In our new paper, @JoshAEngels, @bilalchughtai_, and I find that yes, models finetuned on docs describing a CoT monitor evade detection far more often than unaware models 🧵
2
11
67
9,181
why i changed my mind on what to do with completed tasks: bilalchughtai.substack.com/p…
10
4,037
why i run + some lessons from the past year on the road bilalchughtai.substack.com/p…
8
3,006
bilal 🔶 retweeted
Our new @GoogleDeepMind paper studies novel activation probe architectures for classifying real-world misuse risks. Our research has informed live deployments of probes in Gemini. 🧵
16
58
736
158,155
bilal 🔶 retweeted
NEW PAPER from UK AISI Model Transparency team: Could we catch AI models that hide their capabilities? We ran an auditing game to find out. The red team built sandbagging models. The blue team tried to catch them. The red team won. Why? 🧵1/17
2
22
111
26,772
bilal 🔶 retweeted
Research related to detecting AI deception has a bunch of footguns. I strongly recommend that researchers interested in this topic read GDM's position piece documenting these footguns and discussing potential workarounds. More reactions in 🧵
can we detect when an AI system is being strategically deceptive without relying on behavioural evidence? we, on the GDM mechanistic interpretability team, spent many months working on this problem in our new position piece, we share what we learned:
3
5
38
3,437
can we detect when an AI system is being strategically deceptive without relying on behavioural evidence? we, on the GDM mechanistic interpretability team, spent many months working on this problem in our new position piece, we share what we learned:
5
14
153
35,369
finally, we discuss a further problem that might pose a roadblock to building deception detectors even if the problem of label assignment is solved; namely, that there may not exist a universal set of mechanisms that produce or pinpoint deception in LLMs.
1
23
1,010