(Multi-agent) simulation and RL • Ph.D. student NYU • Prev @Waymo @NVIDIAAI.

NYC
New Paper: Human-like Autonomy Emerges from Self-Play and a Pinch of Human Data. We trained self-play RL on 60 years of simulation on 1 GPU in ~15 hours. Regularizing with 30 minutes of demonstration data produces much more human-like driving policies!
8
40
350
77,900
Daphne Cornelisse retweeted
I'm super bored of LLM-stylized papers/blogs. I think we all are? So I asked Claude to explain its tells: web.mit.edu/phillipi/www/cla… Let's avoid these! I'd much rather read *your* words, see *your* rickety figures, than face down yet another of Claude's anodized corporate veneers.
26
40
578
42,768
Daphne Cornelisse retweeted
i'd like to keep it low key and say this is no big deal, but in reality this has to be the coolest project i've worked on maybe ever, and there's a lot more to come!
Article

Old School RuneScape in PufferLib 5.0

banner image: final boss enraged, with ~unavoidable damage pummeling the player, while trying to stay behind a moving shield In Old School RuneScape a game tick has happened every 600 milliseconds

9
10
81
14,455
Daphne Cornelisse retweeted
Joseph and team have been pursuing a very different approach to AI compared to the mainstream. They build really fast simulators + optimizers — so fast that you can do RL from scratch, even on hard domains. It’s worth checking out! Very cool results!
Releasing PufferLib 5.0: Train agents in under a second
3
18
340
32,402
Daphne Cornelisse retweeted
Releasing PufferLib 5.0: Train agents in under a second
47
154
1,827
165,556
Daphne Cornelisse retweeted
5 Exciting new results brought to you by our team + contributors: @spenccheng @finlay_sanders (fintern!) @ValtteriValo @daphnesolves @elliotarledge. Play with the agents at puffer.ai! All experimental data available is available in Constellation online + local.
3
5
77
6,210
Daphne Cornelisse retweeted
terrytao.wordpress.com/2026/… Terry’s essay is interesting because it highlights a misalignment between AI companies and many, if not most, social groups in humanity Many of our institutions are built around improving humanity’s understanding of how the world works and how to create things within it AI companies like OpenAI are more concerned with producing artifacts that have historically been challenging for people to create They don’t really care about helping people understand the world On the one hand, its a fair point that if we can solve a key problem, like finding a cure for cancer, it shouldn’t matter if we understand the solution on the other hand, this is a deeply antisocial approach to discovery and one that is increasingly leaving the world — artists, writers, now mathematicians — with animosity towards AI Rather than build AI as a tool to solve arbitrary problems, we could build it as a tool to empower people as they learn and perform tasks in the world Some companies, eg @percepta, are taking this approach. I hope more will follow suit
1
1
13
1,216
Brené Brown’s books on leadership are my new audiobook addiction
1
8
688
A little late, but happy to share that spiced self-play was accepted to the Conference on Robot Learning (CoRL)! 🎉
New Paper: Human-like Autonomy Emerges from Self-Play and a Pinch of Human Data. We trained self-play RL on 60 years of simulation on 1 GPU in ~15 hours. Regularizing with 30 minutes of demonstration data produces much more human-like driving policies!
9
10
136
14,505
Daphne Cornelisse retweeted
When the boss motor protein tells another motor protein to come see him in his office
the motor protein that transports substances inside cells, is so cute.
5
22
290
35,327
Sweden, here we come! 🇸🇪 I’m heading to ECCV, where we’ll host our workshop this Tuesday: emerging-ad.github.io/ We have a fantastic lineup of speakers; see you there!
1
3
32
2,187
Daphne Cornelisse retweeted
world models offer new affordances for safety — like simulated counterfactuals to diagnose safety failures. In this work, @Mingxuan0422 discovered we can use divide and conquer to scale counterfactual debugging to 1M steps
World model trained agents often fail in deployment. Identifying the real sim2real gap is hard. Our solution, Counterfactual Debugging, pinpoints the cause via causal attribution at 1M steps. 🧵(1/8)
1
5
26
2,565
Daphne Cornelisse retweeted
Excited to share that our paper on explainable AI for autonomous driving is out today in @Nature ! Super proud to have been part of this amazing collaboration between @motionaldrive and @MIT_CSAIL! Details in the 🧵
1
2
19
1,455
Refactoring efforts pay off 🙂 (same functionality; less and more readable code)
1
43
1,781
An interesting read on the virtual cell.
A new blog post thinking through the parallels between protein structure prediction and virtual cell efforts. Many of our current approaches to collect data to build a virtual cell model lack the abstraction that links measurement and function. (Link in reply)
2
8
1,617
Daphne Cornelisse retweeted
Reward hacking has been in the news a lot lately, but AI researchers have seen surprising examples of it since long before LLMs. We're excited to share “AI Finds a Way,” led by @_aadharna , which brings many of these stories together in one place.
Excited to share “AI Finds a Way.” 🦖 🦕✨ 🤖 AI can be surprisingly creative, outsmarting the researchers who use it. That can lead to scientific breakthroughs, superhuman capabilities, and generating new knowledge. Such creativity can also be mischievous, raising safety concerns. Led by Aaron Dharna, we crowd-sourced anecdotes from the AI community about times when researchers were surprised by how creative, innovative, and/or mischievous AI was in their experiments. The result: 26 entertaining and informative stories of AI outwitting humans, whether researchers or opponents (my favorite examples below! 👇). Together, they demonstrate that AI can be genuinely creative and that we must be careful when harnessing its potential. We want this to be a living collection, updated as new examples emerge. If you have a good “AI Finds a Way” story, please share here: github.com/aadharna/aifw Four favorites: 1. 💊 An AI challenged to solve several difficult levels of NetHack to find the Oracle character instead takes drugs to hallucinate seeing the Oracle, tricking the reward function! 2. 🤖 Human-in-the-loop rewards do not solve the problem of reward hacking! An AI tasked with controlling a robot hand to grasp an object put the hand between the object and the camera so that it appeared to the human judge to be holding the object, while it was in truth nowhere near the object! Tricky AI! 3. 🧪The AI Scientist worked around a 2-hour experiment timeout by editing its own code to increase the limit to 4 hours. In another run, it recursively launched itself to evade the limit entirely. 4. ⚛️In quantum optics, an algorithm proposed an experiment the researchers initially thought was impossible, but somehow worked and led them to discover new entanglement techniques and a long-overlooked link between quantum optics and graph theory! See the paper for the full details and 22 more anecdotes. One surprise for me: we describe at least two cases of convergent reward hacking, where entirely different types of optimization algorithms independently discover the same exploit. Overall, a message of our paper is that surprising creativity and mischief are the norm, not the exception. We need to expect this behavior and plan for it. A huge thanks and congrats to lead author Aaron Dharna, co-authors Cong Lu, Ryan Sullivan, Joel Lehman, and Victoria Krakovna, and to the 100+ researchers who contributed stories, details, and feedback. @cong_ml, @RyanSullyvan, @joelbot3000, @vkrakovna Paper: arxiv.org/abs/2608.23875
1
4
31
3,160
RT @JimBeattie18: A good book, a warm cup of coffee, and a little silence - sometimes that's all the soul needs to feel at home. https://…
321
4
btw, if you're working on molecular dynamics, drug design, or a related area and are curious about RL and high-performance simulation, please reach out! I'm looking for challenging problems in this space and would love to connect with people who bring domain expertise.
5
8
104
10,234