Postdoc @UBC_CS with @jeffclune (RL, Curriculum Learning, Open-Endedness) | PhD from @UofMaryland | Previously RL @SonyAI_global and RLHF @Google

Ryan Sullivan retweeted
📢 3 days until the Continually Self-Improving Robots workshop deadline of Sep 28th (AOE). Working on robots that autonomously learn, adapt, and improve over time? Submit and join us in Austin! csircorl.github.io/website/#…
Today’s robots are impressive. But making them reliable for months across the long tail of the real world is a different challenge. That’s why we’re organizing the Continually Self-Improving Robots workshop at #CoRL2026 csircorl.github.io/website/
5
22
5,168
While most of our field's attention turned toward LLMs, Joseph has been improving tabula rasa RL at an incredible pace. If you're doing RL research I recommend you check out PufferLib. It's ridiculously fast, easy to use, and has a growing community of researchers supporting it.
Releasing PufferLib 5.0: Train agents in under a second
1
1
36
2,939
With this latest release, they have strong results on extremely challenging environment environments like NetHack and Craftax. In fact, they have so many results that some of my favorites didn't make it into the blog post! You can join the discord to see their latest progress.
1
9
130
Ryan Sullivan retweeted
Turning to open-ended innovation as the next frontier for AI makes a lot of sense, I commend @arcprize for moving in this direction, but I’m very curious how they will conceive a “benchmark” for “open-endedness,” two words that seem almost antithetical to each other. In fact, one likely reason that open-ended innovation has lagged behind other areas of AI is how fundamentally resistant it is to benchmarking. Now that doesn’t necessarily mean there’s no hope for an imaginative approach. Attempts at measuring open-endedness go back to Bedau’s activity statistics in the field of artificial life, Several colleagues and I later introduced a measure called “ANNECS — Accumulated Number of Novel Environments Created and Solved” in our paper on Enhanced POET. That’s not an exhaustive list. But there’s never been the kind of benchmark where you can just easily put any systems seamlessly head to head, and there are enormous pitfalls if you get wrong. After all, if “open-ended innovation” ends up equated to “solving a prescribed.hard problem in a creative way” then you risk actually rewarding the opposite of open-endedness, which needs to account for the fact that deciding the “problem” or objective is part of the job of the open-ended system itself. And also, perhaps even more prohibitively for benchmarking, that a key aspect of open-endedness is to be intelligent when you don’t have a defined objective or problem at all! How can that be benchmarked? I still think it’s great that ARC Prize is bringing attention to this part of AI space, and I’d be happy to connect and exchange thoughts on how to get it right if that could be useful.
ARC-AGI-4 will be a benchmark for autonomous open-ended innovation. It will continue our commitment to open-source, giving the research community a shared target for progress that benefits all of humanity. Despite rapid model progress, humans still significantly outperform AI at open-ended invention. This is the meta-skill that unlocks progress across every field of technology. Advanced AI capable of scientific innovation will lead to tremendous new technology, knowledge, and understanding. This is a positive-sum future. We are deeply committed to advancing it. Open source is the foundation for that progress. The knowledge behind frontier AI, not just the technology itself, should be broadly distributed among researchers, academics, and organizations. Any coordinated effort by the AI industry to reduce openness or concentrate access to frontier AI would undermine that positive-sum future. We are committed to advancing a future where everyone can contribute to and benefit from AI progress.
11
22
165
10,566
Ryan Sullivan retweeted
Less than a month ago @RL_Conference I presented my position that general videogame AI is an open-ended problem, but now GPT-6 Astra can apparently play at least half the games I talked about (Celeste, Minecraft, Zork) and many more - so does this change my opinion?
1
4
65
3,678
Ryan Sullivan retweeted
1/ What do RL, LLMs and evolutionary algorithms have in common? Give them an objective and they find a loophole - like a Pokemon agent rewarded for exploring that learned to stand still and watch animated flowers. AI Finds A Way collects 26 such cases: arxiv.org/abs/2608.23875
Excited to share “AI Finds a Way.” 🦖 🦕✨ 🤖 AI can be surprisingly creative, outsmarting the researchers who use it. That can lead to scientific breakthroughs, superhuman capabilities, and generating new knowledge. Such creativity can also be mischievous, raising safety concerns. Led by Aaron Dharna, we crowd-sourced anecdotes from the AI community about times when researchers were surprised by how creative, innovative, and/or mischievous AI was in their experiments. The result: 26 entertaining and informative stories of AI outwitting humans, whether researchers or opponents (my favorite examples below! 👇). Together, they demonstrate that AI can be genuinely creative and that we must be careful when harnessing its potential. We want this to be a living collection, updated as new examples emerge. If you have a good “AI Finds a Way” story, please share here: github.com/aadharna/aifw Four favorites: 1. 💊 An AI challenged to solve several difficult levels of NetHack to find the Oracle character instead takes drugs to hallucinate seeing the Oracle, tricking the reward function! 2. 🤖 Human-in-the-loop rewards do not solve the problem of reward hacking! An AI tasked with controlling a robot hand to grasp an object put the hand between the object and the camera so that it appeared to the human judge to be holding the object, while it was in truth nowhere near the object! Tricky AI! 3. 🧪The AI Scientist worked around a 2-hour experiment timeout by editing its own code to increase the limit to 4 hours. In another run, it recursively launched itself to evade the limit entirely. 4. ⚛️In quantum optics, an algorithm proposed an experiment the researchers initially thought was impossible, but somehow worked and led them to discover new entanglement techniques and a long-overlooked link between quantum optics and graph theory! See the paper for the full details and 22 more anecdotes. One surprise for me: we describe at least two cases of convergent reward hacking, where entirely different types of optimization algorithms independently discover the same exploit. Overall, a message of our paper is that surprising creativity and mischief are the norm, not the exception. We need to expect this behavior and plan for it. A huge thanks and congrats to lead author Aaron Dharna, co-authors Cong Lu, Ryan Sullivan, Joel Lehman, and Victoria Krakovna, and to the 100+ researchers who contributed stories, details, and feedback. @cong_ml, @RyanSullyvan, @joelbot3000, @vkrakovna Paper: arxiv.org/abs/2608.23875
2
9
34
4,752
Ryan Sullivan retweeted
It seems like every week now there is another story of an AI model making new scientific discoveries, and every week there is another story of models breaking out of sandboxes or finding some bizarre way around the rules. Our new paper: “AI Finds A Way” collects these stories!
Excited to share “AI Finds a Way.” 🦖 🦕✨ 🤖 AI can be surprisingly creative, outsmarting the researchers who use it. That can lead to scientific breakthroughs, superhuman capabilities, and generating new knowledge. Such creativity can also be mischievous, raising safety concerns. Led by Aaron Dharna, we crowd-sourced anecdotes from the AI community about times when researchers were surprised by how creative, innovative, and/or mischievous AI was in their experiments. The result: 26 entertaining and informative stories of AI outwitting humans, whether researchers or opponents (my favorite examples below! 👇). Together, they demonstrate that AI can be genuinely creative and that we must be careful when harnessing its potential. We want this to be a living collection, updated as new examples emerge. If you have a good “AI Finds a Way” story, please share here: github.com/aadharna/aifw Four favorites: 1. 💊 An AI challenged to solve several difficult levels of NetHack to find the Oracle character instead takes drugs to hallucinate seeing the Oracle, tricking the reward function! 2. 🤖 Human-in-the-loop rewards do not solve the problem of reward hacking! An AI tasked with controlling a robot hand to grasp an object put the hand between the object and the camera so that it appeared to the human judge to be holding the object, while it was in truth nowhere near the object! Tricky AI! 3. 🧪The AI Scientist worked around a 2-hour experiment timeout by editing its own code to increase the limit to 4 hours. In another run, it recursively launched itself to evade the limit entirely. 4. ⚛️In quantum optics, an algorithm proposed an experiment the researchers initially thought was impossible, but somehow worked and led them to discover new entanglement techniques and a long-overlooked link between quantum optics and graph theory! See the paper for the full details and 22 more anecdotes. One surprise for me: we describe at least two cases of convergent reward hacking, where entirely different types of optimization algorithms independently discover the same exploit. Overall, a message of our paper is that surprising creativity and mischief are the norm, not the exception. We need to expect this behavior and plan for it. A huge thanks and congrats to lead author Aaron Dharna, co-authors Cong Lu, Ryan Sullivan, Joel Lehman, and Victoria Krakovna, and to the 100+ researchers who contributed stories, details, and feedback. @cong_ml, @RyanSullyvan, @joelbot3000, @vkrakovna Paper: arxiv.org/abs/2608.23875
1
7
23
2,549
Reward hacking has been in the news a lot lately, but AI researchers have seen surprising examples of it since long before LLMs. We're excited to share “AI Finds a Way,” led by @_aadharna , which brings many of these stories together in one place.
Excited to share “AI Finds a Way.” 🦖 🦕✨ 🤖 AI can be surprisingly creative, outsmarting the researchers who use it. That can lead to scientific breakthroughs, superhuman capabilities, and generating new knowledge. Such creativity can also be mischievous, raising safety concerns. Led by Aaron Dharna, we crowd-sourced anecdotes from the AI community about times when researchers were surprised by how creative, innovative, and/or mischievous AI was in their experiments. The result: 26 entertaining and informative stories of AI outwitting humans, whether researchers or opponents (my favorite examples below! 👇). Together, they demonstrate that AI can be genuinely creative and that we must be careful when harnessing its potential. We want this to be a living collection, updated as new examples emerge. If you have a good “AI Finds a Way” story, please share here: github.com/aadharna/aifw Four favorites: 1. 💊 An AI challenged to solve several difficult levels of NetHack to find the Oracle character instead takes drugs to hallucinate seeing the Oracle, tricking the reward function! 2. 🤖 Human-in-the-loop rewards do not solve the problem of reward hacking! An AI tasked with controlling a robot hand to grasp an object put the hand between the object and the camera so that it appeared to the human judge to be holding the object, while it was in truth nowhere near the object! Tricky AI! 3. 🧪The AI Scientist worked around a 2-hour experiment timeout by editing its own code to increase the limit to 4 hours. In another run, it recursively launched itself to evade the limit entirely. 4. ⚛️In quantum optics, an algorithm proposed an experiment the researchers initially thought was impossible, but somehow worked and led them to discover new entanglement techniques and a long-overlooked link between quantum optics and graph theory! See the paper for the full details and 22 more anecdotes. One surprise for me: we describe at least two cases of convergent reward hacking, where entirely different types of optimization algorithms independently discover the same exploit. Overall, a message of our paper is that surprising creativity and mischief are the norm, not the exception. We need to expect this behavior and plan for it. A huge thanks and congrats to lead author Aaron Dharna, co-authors Cong Lu, Ryan Sullivan, Joel Lehman, and Victoria Krakovna, and to the 100+ researchers who contributed stories, details, and feedback. @cong_ml, @RyanSullyvan, @joelbot3000, @vkrakovna Paper: arxiv.org/abs/2608.23875
1
4
31
3,161
We wrote this for a broad audience, so no technical background is required. If you’re curious about the history of reward hacking, or just looking for entertaining stories of machines doing things researchers didn't expect, there should be something here for you.
1
4
71
I’m at RLC presenting a workshop paper on Meta-Exploration with @BenNorman451 (starting now at poster 17!) Looking forward to meeting new people and chatting about exploration, curriculum learning, and open-ended environments! Feel free to dm me if you want to meet up!
1
1
28
2,453
If you want to really understand policy gradients beyond common knowledge, you should follow Chris. I learned a lot from his papers at the start of my PhD.
Most people think REINFORCE is the simplest policy gradient algorithm. But you can actually go a step simpler and multiply all of the log probabilities by the whole-trajectory return and never compute the per timestep return-to-go. I wonder if there any cases where this could be genuinely useful? Follow for more useless RL trivia 🙂
2
4
1,582
Ryan Sullivan retweeted
New Paper: Human-like Autonomy Emerges from Self-Play and a Pinch of Human Data. We trained self-play RL on 60 years of simulation on 1 GPU in ~15 hours. Regularizing with 30 minutes of demonstration data produces much more human-like driving policies!
8
40
350
77,945
Ryan Sullivan retweeted
We never really knew how to train nonlinear RNNs well… BPTT struggled with vanishing grads (no long-range memory) and sequential rollout (hard to parallelizable). What if instead an oracle told us the optimal memory state m_t at each step? Then the RNN could do one-step supervised learning on (m_t, x_{t+1}) → m_{t+1} labels. We call this Supervised Memory Training (SMT): a replacement for BPTT that trains RNNs without unrolling them. SMT is time-parallelizable and solves vanishing gradients. Website: akarshkumar.com/smt/ arXiv: arxiv.org/abs/2606.06479
18
119
807
187,815
I’m in SF for the week if anyone wants to grab a coffee and chat about research! I’ve been thinking a lot about exploration, search, and open-ended environments recently. Feel free to DM me!
2
2
21
3,556
Ryan Sullivan retweeted
A new and possibly controversial perspective: In this video, I explain the sense in which generative AI trained by supervised learning is incapable of making novel discoveries. piped.video/K5LAFEjTlBA The text of the speech: AI Creativity and Discovery Good day ladies and gentlemen. I regret that I am unable to be with you all today to engage in a back-and-forth discussion, but I am nevertheless pleased to be able to share with you, via this recording, some high-level thoughts about the current and future state of artificial intelligence, and in particular about AI’s relationship to science and mathematics, which is, as I understand it, the central focus of this meeting and of the SAIR Foundation. I would like to start with an old joke; I am sure you have heard it before. It is the one about the researcher whose work is being evaluated, and the review comes back, and says “This work is both novel and good. Unfortunately, the parts that are good are not novel, and the parts that are novel are not good.” My first point about AI is that this assessment applies exactly to large parts of AI as we know it today. Not all of today’s AI, but a large part of it. Pretty much all of what we mean by “Generative AI”---which includes large language models, and the images and video models, and even the new methods for learning world models. All of these AIs take large numbers of examples and produce a “model” which behaves similar to the examples, that is, which generates text like people, or images like artists or nature, and videos like we find on the internet. Don’t get me wrong, Generative AI can be extremely useful. No doubt about that. But the assessment of the joke still applies. These systems can produce output that is both novel and good, but not at the same time. In many ways this is just absolutely not a problem. When we ask an AI for an answer from the internet, or to summarize a document, we don’t want it to be novel. We are happy if the quality of the answer, the goodness, comes from the source material—from the people who wrote the document or the articles on the internet. If the AI’s answer is novel it means it is going beyond the source material, adding something beyond it. This is what we call “hallucinations”. In most cases, we don’t like it when the AI makes something up, when it adds something novel. One exception, of course, is when we are looking not for facts or reality, but for fiction and entertainment. We might ask for a bedtime story for a child, or an image based on existing images on the internet but which is nevertheless different and distinct from them. In these cases, it is never easy for us to know how creative the AI is actually being, as we do not know how close the AI’s story, poem, or image is to the source material. In a real practical sense we can not know this because the internet is too big, the possible sources that the AI may draw upon are too numerous. When we ask for a fiction or novelty, the AI can give it to us because its processing is in part stochastic. Every decision can go multiple ways and will go different ways and produce a different trajectory every time. The trajectory can be random—and thus novel—or it can be based on the training data—and thus “good” because the training data is good, sourced from people or reality. Thus, the trajectory is either novel or good—based on randomness or based on data—but never both at the same time. Really, I think it is okay if the output of Generative AI is never good and novel at the same time. For the researcher in the joke this is a devastating criticism, but for most things it is not, and for Generative AI it is not. Generative AI is meant to be a mimic. This is what supervised learning is for. Generative AI can be extremely useful, even when it just mimics, if it is faster, or cheaper, or smaller, or more customizable, or more copy-able, than the thing being mimicked. It is okay if Generative AI cannot be both novel and good at the same time. It is still a transformative technology. But it is a limitation. And remember we are here to use AI for science and mathematics, and for these areas the assessment of the reviewer in the joke is devastating. For these areas we need true creativity and discovery. Generative AI—or Mimicking AI—will never get where us there. For these we need something more, and indeed we have something more in other parts of AI. We have many AI systems which can give us more. We have AlphaGo with its world-changing move 37, or AlphaZero with its brilliant original chess-playing style. We have GT-Sophy that drives simulated racecars better than any human. We have AlphaFold and AlphaProof and Claude-Code, which have brought true advances in science, mathematics, and programming. We have RL-Lyft which optimizes the assignment of cars to passengers in the ride-hailing business. All these systems have found things that are both novel and good. And, truth be told, some language models have been augmented in ways that make them more than Generative AI based on supervised learning. All these systems have some additional features that make them capable of true creativity and true discovery. It is important for us to recognize what this is—and that it is not present in ordinary, garden-variety Generative AI. It is something that can not come from just supervised learning, from learning from examples. What is it? Well, it is a simple thing, a commonsense thing. It is not new. We have many names for it, but unfortunately none of them are very good names. I will call it Discovery. Basically, Discovery is just the idea of trying many things and seeing which of them work, then keeping those that worked the best. Evolution by natural selection works this way. The scientific method works this way. And just ordinary life and learning works this way. We try things and remember what works. What could be more obvious? In this behavioral case, psychology has two names for it— “instrumental learning” and “operant conditioning”—and in machine learning it is what we mean by “reinforcement learning”. We also see the idea of Discovery in planning and combinatorial search—anything that involves the idea of “generate and test”. The essence of Discovery is to combine three steps: 1. Variation, 2. Evaluation, and 3. Selective retention. Of course, I am not the first to say this. I am not the first to point out that this combination of steps is key to science, to evolution by natural selection, and to animal behavior. I think particularly of papers by Donald Campbell, by Daniel Dennett, and by Gary Cziko. What is new in my remarks is to directly relate the idea of Discovery to modern AI to help us see that it is not present in supervised learning or Generative AI—in particular, that Discovery is not present in backpropagation or gradient descent. Let me say explicitly what is missing from Generative AI. As we have remarked, these systems do have a stochastic aspect, so they do generate a variety of trajectories and behavior. What is missing is the Evaluation step. The generator was pre-trained by supervised learning, leaving no way at runtime to Evaluate what it generates. And of course without Evaluation there can be no Selective retention, and thus no Discovery. The variation can bring novelty, but without evaluation there is no Discovery, and arguably, no creativity. That is, I would say that creativity requires that the new things generated be Evaluated. Without evaluation, and retention of the best, there is nothing created. The novelty flickers into existence but, if its value is unrecognized, it flickers away and is lost. In many cases, Evaluation is done by people to make a discovery. As when we have Generative AI make many pictures for us, and then we pick the one that we like the best. The human+AI system completes the discovery. In many other cases, the Evaluation comes from a clear objective. Some moves lead to checkmate, some steps lead to a proof, some actions result in high reward, some genotypes make more copies, some theories explain the data better. Some prefer the Variation step to be called Blind variation, where “blind” here means that it is uninformed, a shot in the dark. It does not need to be completely uninformed; a good scientist does not select theories to test at random. But neither can it be completely informed and determined. There must be some uncertainty about where the answer lies in order for there to be a discovery. In practice, the variation is partly informed and partly blind, but it is the blind part that corresponds to the discovery. Now let us briefly go all the way to modern deep learning, to the backpropagation algorithm. At first it might seem that backpropagation is incapable of discovery because it is deterministic and thus incapable of variation. But this is not correct. The weight updates of backprop are deterministic, but the weights are initialized to small random values. The random initialization is often downplayed, but in fact it is a necessary form of variation; it must be done properly to get good performance. In backprop this Variation is done once, at network initialization, so its effect is temporary, and later the network may lose its ability to learn. This is the weakness of deep learning that is alleviated with a new algorithm that my group presented in Nature a couple of years ago. Our “continual backpropagation” made one small change: every so often a less-used neuron would be re-initialized to small random weights. This allows the variation to continue and plasticity to be retained. Although there is much more to be said about Creativity and Discovery, this is the key point: they are more than supervised learning, more than pattern recognition, more than prediction, and more than world modeling. Those things are important, but they alone will not bring us to discovery. Discovery requires Evaluation from a person or from an explicit goal, and only in the latter case will we attain full autonomy. So that is my call to arms. If we want the full power of AI scientists, then we should share the goals with them so they can create, evaluate, discover, and in these ways fully participate in achieving the goals. Let’s be bold! Let’s fully automate Creativity and Discovery!
104
284
1,688
695,713
It’s really incredible to see a company fully dedicated to open-endedness at this scale. Congrats to the team, I’m looking forward to seeing what you create!
Thrilled to share that we founded Recursive to create AI that safely conducts experiments on how to improve itself in an open-ended process of endless, automated scientific discovery. As I wrote in my 2019 AI-generating algorithms paper, this will likely be the fastest path to superintelligence. Our work since has shown the power of this approach. Excited to scale up and improve upon ideas like the Darwin Gödel Machine, HyperAgents, ADAS, OMNI, ALMA, The AI Scientist, PromptBreeder, Rainbow Teaming, Automated Capability Discovery, and other work on open-ended and AI-generating algorithms. We’ve assembled a dream team of researchers and significant resources to pursue this vision. My amazing co-founders are pictured here, and we have an all-star team of founding members (we’re over 25 and growing). Please join us if you are interested! Follow our progress @Recursive_SI
2
1
22
2,933