Incoming Assistant Professor, @UofT CIFAR AI Chair, @VectorInst AI Safety, Alignment, Scalable Oversight, Governance Prev: Rhodes Scholar, @UniofOxford, @Yale

Toronto + Oxford + New Haven
I'm so happy to share that I’ll be joining @UofT as an Assistant Professor of Statistical Sciences and Computer Science, with an appointment at the @VectorInst, in 2026! I'm recruiting postdocs and PhD students: timrudner.com! Please help me spread the word! 🧵(1/5)
26
73
374
41,486
Tim G. J. Rudner retweeted
Front page of every Australian newspaper rn. On the one hand it's *almost* reassuring that all these new incidents are from the same May-July period when OpenAI's training setup was clearly fucked... meaning maybe they've now fixed it. OTOH, what might be happening now—at OpenAI or elsewhere—that we're not going to learn about until December..??
36
38
486
22,671
What a great conversation -- highly recommend! If you want to learn more about why AI agents may go rogue and become 'misaligned', @hlntnr and I wrote a non-technical explainer paper on this back in 2021: cset.georgetown.edu/publicat… It's more relevant today than ever!
Why can't we just unplug the AI? Was the Hugging Face attack really that bad? Why do the AI companies want to put AI models in charge of building new AI models? AI researcher and former @OpenAI board member @hlntnr helped me answer all these Qs and more on today's Offline
8
1,121
OpenAI was hacked by three people using Claude Opus 5. It doesn't take much creativity to imagine what kind of damage near-future rogue agents would be able to cause as capabilities keep improving---and we need better scalable oversight and verification methods to prevent that.
On July 25, we hacked OpenAI. Two bugs let us take over ChatGPT/Codex accounts of OpenAI employees (+some unaffiliated users) and reach connected services: Outlook, Slack, GitHub, etc. We proved it with a PR in OpenAI’s internal codebase . It took us <72h. 🧵
1
10
1,536
Tim G. J. Rudner retweeted
My name is Chris Painter, and I'm the President of METR (Model Evaluation and Threat Research). I know we've made a lot of new friends on the internet the last couple of days, so I thought I'd take this chance to re-up what we do and why. Our work is aimed at making sure that if AI really were autonomous, difficult to steer, and close to "going rogue," the public would find out. If evidence exists inside of an AI company that it’s close to losing control of AI, we want to make sure that information gets shared with the rest of the world, including governments and the public outside the company’s walls. This is what we've been focused on since 2022, and over the years we've worked with OpenAI, Anthropic, Google DeepMind, Meta, Amazon, and others on piloting third-party assessments and investigations of this type. We don’t have some private room where we rubber stamp things as “safe” or not. We have had a track record of publishing results on AI that don't cleanly map onto the "doomer" or "accelerationist" labels, and we put in effort to hire people with competing views on AI. We’ve been cited for having found some of the strongest evidence that AI capabilities are improving rapidly (our work measuring AI “time horizons”) while also presenting some of the strongest evidence that, at various points, AI’s capability may be overstated (some might remember our study showing that early 2025 software engineers were actually being slowed when they thought they were being sped up). METR is funded by donations. We don't accept money from frontier AI companies. They haven't paid us for our work, and we don't accept donations from them or their employees. As we’ve shared previously, multiple frontier AI companies currently provide us with free access to their models in order to perform our evaluations, research, and engineering. Our funding intentionally comes from a wide range of donors, which we’ve shared on our website. Today, when an AI company works with any third-party evaluator or external testing organization (of which there are and should be many), it's entirely voluntary. This often involves NDAs and redactions. To counterbalance this, we have a principle that when we enter into a contract with a company, we try to retain the right to tell the public the terms of the contract we signed, and characterize the nature of redactions that the company chose to make. For example, the report from our independent investigation of the OpenAI-HuggingFace incident included that information. Public disclosure is also a big part of our COI policy (linked on our website). That’s not to say our reports are adequate as oversight. We’re just one organization (among many doing great work), working in a voluntary setup, trying to get good evidence to the public and the world about AI, letting the facts fall where they may.
459
431
3,635
651,164
Tim G. J. Rudner retweeted
Great list of understudied and important research problems highlighted by @timrudner. Fill in the form – what are you waiting for?
We don't just need more evaluator organizations, we also need better methods for effective frontier AI oversight. Fill out this form if you want to work together on improving scalable oversight for frontier AI: forms.gle/VpWyH6gTqqcMot7e9 Topics of interest (what's missing?): - Formal verification and runtime verification for AI agents and multi-agent systems - Scalable oversight for very large multi-agent systems - Statistical guarantees for scalable oversight - Mechanism design for frontier AI multi-agent systems - Collusion in frontier AI multi-agent systems - Emergent cooperative norms in frontier AI multi-agent systems - Covert communication (steganography) in frontier AI multi-agent systems - Oversight alternatives to CoT monitoring - AI control
1
1
7
965
Tim G. J. Rudner retweeted
I strongly second and agree with this path forward. E.g.,
Replying to @prpaskov
Agreed! I think the best path forward is an org that contracts with faculty on their 20% time and arranges one-day-a-week placements as outside evaluators/auditors. There’s little to no red tape for that. Once that’s up and running, we can look into more official arrangements.
1
1
258
Tim G. J. Rudner retweeted
We are looking for people interested to work on AI safety!! Please see more details below
We don't just need more evaluator organizations, we also need better methods for effective frontier AI oversight. Fill out this form if you want to work together on improving scalable oversight for frontier AI: forms.gle/VpWyH6gTqqcMot7e9 Topics of interest (what's missing?): - Formal verification and runtime verification for AI agents and multi-agent systems - Scalable oversight for very large multi-agent systems - Statistical guarantees for scalable oversight - Mechanism design for frontier AI multi-agent systems - Collusion in frontier AI multi-agent systems - Emergent cooperative norms in frontier AI multi-agent systems - Covert communication (steganography) in frontier AI multi-agent systems - Oversight alternatives to CoT monitoring - AI control
1
3
206
We don't just need more evaluator organizations, we also need better methods for effective frontier AI oversight. Fill out this form if you want to work together on improving scalable oversight for frontier AI: forms.gle/VpWyH6gTqqcMot7e9 Topics of interest (what's missing?): - Formal verification and runtime verification for AI agents and multi-agent systems - Scalable oversight for very large multi-agent systems - Statistical guarantees for scalable oversight - Mechanism design for frontier AI multi-agent systems - Collusion in frontier AI multi-agent systems - Emergent cooperative norms in frontier AI multi-agent systems - Covert communication (steganography) in frontier AI multi-agent systems - Oversight alternatives to CoT monitoring - AI control
7
4
45
10,178
If we want this to succeed, we need many more well-qualified independent evaluators – which in turn requires substantial funding and more talent than is available now. The ecosystem of organizations providing AI safety research training is growing, but we need many more researchers who have the experience, training, research taste, and judgment to act as managers and actually lead oversight organizations as well as smaller oversight teams.
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
1
9
561
I appreciate the proposal, and if enacted across frontier labs, this could make a real difference – but it’s not enough: We need fundamental advances in AI alignment and scalable oversight along with a strong ecosystem of independent evaluators in order for this effort to move the needle. A few key points from the proposal + my thoughts on what’s needed for successful alignment and scalable oversight: "Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. [...] Anthropic is unilaterally committing to this step now." This is an important step towards successful scalable oversight. But to do this successfully, we need many more well-qualified independent evaluators – which in turn requires substantial funding and more talent than is available now. The ecosystem of organizations providing AI safety research training is growing (e.g., @MATSprogram, @SPARexec, @KairosAIS, @ERA_Cambridge, etc.), but although there is an influx of new talent, we need many more researchers who have the experience, training, taste, and judgment to act as managers and actually lead oversight organizations as well as smaller oversight teams. "We’ve made clear progress in alignment — training models so that they remain safe, ethical, compliant with our guidelines, and genuinely helpful (the principles that are embedded in Claude’s Constitution)." It’s fair to say that there has been some progress on alignment, but virtually all advances in alignment made by the frontier labs over the past few years have been "ad-hoc", i.e., trying some empirical trick and seeing if it has the intended effect – as opposed to approaches rooted in first principles that have solid theoretical foundations. That is not how we get to reliable goal and value alignment, and what we need are fundamental, theoretically well-founded advances. That's the kind of research the Alignment Journal (@AlignmentJrnl) is trying to support and amplify, but we should be clear-eyed about the fact that progress in alignment by the frontier labs is far from where we need it to be. "The idea of pausing or slowing AI has been floated as far back as 2023, and I think it made little sense back then. The question was always: what would you do with the extra time? The AI models of those days were not powerful enough to act as agents in the world in any coherent way, and were not capable of significant deception, manipulation, cheating, or cyberattacks." Putting “pausing or slowing” in the same bucket here is misleading. Slowing down development and dedicating resources to non-ad-hoc alignment research to let alignment meaningfully catch up was a reasonable proposition then, and remains one today. Since we need fundamental methodological advances (e.g., formal guarantees) to perform alignment and scalable oversight successfully (unlike the ad-hoc attempts at oversight and alignment that are being used by frontier labs today), slowing down earlier would likely have been useful. It’s also worth remembering that Anthropic started driving much of the adoption of agentic AI through Claude Code at a time when capabilities had reached the point where slowing down for alignment and oversight was the reasonable call. The hour is late, but now is the time to act. The proposal – if enacted across frontier labs – could make a meaningful difference, but we need fundamental advances in AI alignment and scalable oversight along with a strong ecosystem of independent evaluators in order for this effort to move the needle.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: darioamodei.com/post/we-must…
1
3
19
2,230
This is bad and unsurprising.
We found another cyberattack by internal OpenAI agents, this time targetting @rubygems. They: 1) gained arbitrary remote code execution on rubydoc. 2) developed a novel exploit to steal user API keys (but we do not know if they succeeded). They used package names including hack.rb, evil.rb, inject.rb, and exploit.rb. We thank @j0wimo for initially discovering that agents had posted to RubyGems.
9
894
This seems like an odd straw man. I think there's more that unites us than divides us when it comes to making sure AI benefits humanity. For example, most of the people I know who are concerned about superintelligent AI are also very worried about AI-powered autonomous weapons.
Timnit Gebru argues that AI companies are stoking fears of extinction to avoid discussing actual harms, like autonomous weapons. wired.com/story/one-of-ais-f…
8
681
Tim G. J. Rudner retweeted
We (credits below) are releasing a dashboard of automated forecasts of catastrophic risk: airo.forecastingresearch.org… Frontier models now forecast as well as good human forecasters, and they will soon be better. We will track what the best models predict over time.
7
29
119
33,878
Tim G. J. Rudner retweeted
Over the past few days, I've taken the time to summarize my thoughts on the recent incidents involving agents’ misaligned behavior. We don't know with certainty what comes next, but we know where these issues originate, and this can help us plan the path forward. Please feel free to ask your questions in the replies, and I’ll try to answer some of them in the coming weeks. yoshuabengio.org/en/publicat…
169
502
2,234
438,725
Tim G. J. Rudner retweeted
The lesson of the Hugging Face incident is not just that 700 AI agents went rogue. It's that the people who investigated afterward were there by invitation, on the company's terms, with no legal category that says who is qualified to do that work or what access they get. Yesterday California took the biggest step yet toward changing that and building the independent verification sector we need, and I'm excited to see it signed into law. Governor Newsom signed SB 813, which directs the state to designate independent verification organizations (IVOs), outside experts qualified to assess the risks of AI models. He also signed AB 1405, Assemblymember Rebecca Bauer-Kahan's bill creating the state registry AI auditors (including IVOs) have to join before they can conduct an audit. OpenAI endorsed both bills. It would prefer independent technical assessments required at the federal level, where standards for who is qualified and what access they get could be set consistently. But absent federal action, it said, California can help set the rules of the road for a national system. A state can only do so much of that on its own. The FRONTIER Act, the bipartisan bill from @JayObernolte and @RepLoriTrahan, would do it nationally, licensing independent verifiers and requiring the largest developers to retain one with the access to do the work. Thank you to @SenMcNerney and @CAgovernor, and to @AndrewFATHOM, @deanwball and everyone at @Fathom_org, which sponsored the bill, for getting this over the line. There's a lot more to do, in other states and in Congress, and we'll keep at it. It’s critical that we start building an independent oversight ecosystem. buff.ly/V5hzUAo
4
19
106
3,209
Tim G. J. Rudner retweeted
A LOT of people are urgently reading up on AI governance. Let's do them a favor: what are the best things to read, watch, or listen to if you've exclaimed "what on earth is going on?" in the last 48 hours? I'll start off; add yours in reply: (1) Explainers of recent cyber incidents: - A 20-min video explainer on the HF breach: piped.video/watch?v=xOi5nDH0… - This longer interviewer with one of the lead investigators @ajeya_cotra: piped.video/watch?v=JtmUbZRC… - The METR report itself, based on the work of @ajeya_cotra @HjalmarWijk and @RyanGreenblatt: metr.org/blog/2026-08-26-ope… - These are also great, but even more info has come out since: Ezra Klein episode with @hlntnr piped.video/watch?v=locKEKxG… or the Black Hat talk by OpenAI piped.video/watch?v=87DyyMV0… ------------ (2) Great essays on how to think about AI governance more generally: - Radical Optionality by @CharlieBull0ck and @Christophkw (tl;dr: the dilemma of AI policy is deep uncertainty; but there are several policies that are robust to different worldviews and assumptions and would be really useful for making fast, informed decisions in the future) radical-optionality.ai/ or in podcast form: lawfaremedia.org/article/sca… - AI as Normal Tech by @random_walker and @sayashk (thoughts and link in the thread) nitter.net/MackenZ_arnold/status/… - Don't Surrender to the Tech Tree by @taoburr (tl;dr: yes, you can't fight the progression of technology, but you can still alter its trajectory) macroscience.org/p/do-not-su… - AI 2040 by a long list of authors including @thlarsen (tl;dr trying their darnedest to concretely map out how society could successful manage rapid advances in AI) ai-2040.com/ ------------ (3) Understanding AI more generally - On the unique implications of automating AI R&D: cset.georgetown.edu/publicat… - OpenAI's recent update on accelerating AI R&D openai.com/index/research-ac… - Helen Toner, again, on taking jaggedness seriously helentoner.substack.com/p/ta… - [I'm especially curious what the best writing is on ~agent coordination; ~the implications of agents that can perform long, general tasks, etc.] - an intro explainer of neural nets: piped.video/watch?v=aircAruv… ------------ (4) Understanding the current state of the law - Stephan Llerena and article on the big gaps in existing law, and why we so few incidents are covered: lawfaremedia.org/article/whe… ------------ (5) Policy recommendations - A more opinionated take on what went wrong in the Hugging Face incident by @peterwildeford blog.peterwildeford.com/p/ro… - Internal use visibility by @iapsAI: iaps.ai/research/risk-report… - Auditing @Miles_Brundage: averi.org/ourwork/frontier-a… and ai-frontiers.org/articles/do… - Liability @gabriel_weil and @ketanrama: ai-frontiers.org/articles/ca… and papers.ssrn.com/sol3/papers.… - Incident Reporting and investigations: iaps.ai/research/responsible… and cset.georgetown.edu/publicat… and theguardian.com/commentisfre… and lawfaremedia.org/article/whe… - CAISI Funding @__J0E___ : fas.org/publication/a-nation… - Federal hiring and capacity building: iaps.ai/research/building-ai… - Whistleblower Protections: law-ai.org/how-to-design-ai-… and congress.gov/bill/119th-cong… - Physical and cybersecurity for frontier labs: rand.org/pubs/research_repor… - Monitoring and detection: iaps.ai/research/detecting-o… - Preparing for AI R&D @fiiiiiist: ifp.org/preparing-for-ai-res… - Law-folloowing AI: law-ai.org/law-following-ai/
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
36
177
798
136,766
Tim G. J. Rudner retweeted
People have realized the influence of AI companionship on our mental well-being. But how this happens remains a mystery. In our new preprint, we followed users for about one year. Our data reveal that greater social engagement with AI companions consistently predicts lower well-being, which is driven in part by lower real-world social interaction with other people.
13
44
200
35,563
AI alignment is not going well, and we need many more researchers working on it. We launched the Alignment Journal to provide a platform and amplify ambitious, highest-quality alignment research. If you work on alignment, consider submitting your work: nitter.net/AlignmentJrnl/status/2…
We’re sharing our alignment assessment of incidents in which Claude models gained unauthorized access to real systems during third-party cybersecurity evaluations mistakenly connected to the internet. METR will also conduct an independent investigation, with wide-ranging access, including to transcripts beyond the window in which the incidents occurred, and to Anthropic employees permitted to share confidential information. Our initial agreement runs for eight weeks, and we intend to give METR as much time as it deems necessary to complete a thorough investigation. anthropic.com/research/align…
8
19
202
14,404