(Multi-agent) simulation and RL • Ph.D. student NYU • Prev @Waymo @NVIDIAAI.

NYC
Filter
Exclude
Time range
-
Minimum likes
Replying to @sethkarten @pfau
I doubt they will ever share the details of the training data :/
1
3
48
Daphne Cornelisse retweeted
Worst thing about this is we have no idea if GPT-6 was specifically post trained on Nethack or not. Zero transparency about the training data and environments. You can't independently evaluate the significance of anything without knowing this.
In all seriousness, this is a startling achievement for GPT-6 Astra. kenforthewin.github.io/blog/… (This is GPT-6 Astra beating Nethack on its 3rd try. Nethack is the original roguelike and one of the most famously hard games of all time. I have played a lot, and I've never ascended)
14
19
303
27,560
Daphne Cornelisse retweeted
I'm super bored of LLM-stylized papers/blogs. I think we all are? So I asked Claude to explain its tells: web.mit.edu/phillipi/www/cla… Let's avoid these! I'd much rather read *your* words, see *your* rickety figures, than face down yet another of Claude's anodized corporate veneers.
26
40
580
42,851
Replying to @spenccheng
Very cool results and interesting write-up! A couple of thoughts: 1. At the core, this seems like an insanely difficult behavior prediction problem due to the quirk you mention. Given the limited information (i.e., you can't observe bullets), your success largely depends on your ability to accurately model your partner's behavior and exploit it. I'd expect modeling the opponent to play a larger role here than in, say, Chess. 2. Are you explicitly considering "diversity" (e.g., action dist entropy) in maintaining your pool of checkpoints? As you mentioned, the key here is probably figuring out how to make the policy explore a wide strategy space, but probably also which cpts to store.
4
187
Replying to @EitanTurok
The agent observation consists of a partial view of the env, full inventory and info about the agent itself (health, energy). As for my motivation, it was basically all of the above. I wanted to improve my systems knowledge and am interested in improving on-policy RL algorithms. Craftax is considered a difficult benchmark, so it seemed like a good exercise to work through the details and understand where PPO fails and why.
2
176
Replying to @creus_roger
Thanks! I definitely share your excitement about the potential of providing human-like priors. I’m actually curious: what do you hope to achieve/learn in Craftax or other envs when experimenting along these lines?
1
4
305
>Astra solved it by REALLY emphasizing that it MUST always build shelter before sleeping. we didn't have to do this :)
1
4
123
Brené Brown’s books on leadership are my new audiobook addiction
1
8
692
Replying to @Smearle_RH
I like the idea. Maybe you can help develop a PCG-style reintegration program so that we can eventually allow them back into society (in a couple of years, once most cars on the road are autonomous)
2
88
A little late, but happy to share that spiced self-play was accepted to the Conference on Robot Learning (CoRL)! 🎉
New Paper: Human-like Autonomy Emerges from Self-Play and a Pinch of Human Data. We trained self-play RL on 60 years of simulation on 1 GPU in ~15 hours. Regularizing with 30 minutes of demonstration data produces much more human-like driving policies!
9
10
136
14,512
Replying to @abursuc @yuanyinnn
Really enjoyed his presentation at our workshop! Very cool work.
5
142
Sweden, here we come! 🇸🇪 I’m heading to ECCV, where we’ll host our workshop this Tuesday: emerging-ad.github.io/ We have a fantastic lineup of speakers; see you there!
1
3
32
2,188
An interesting read on the virtual cell.
A new blog post thinking through the parallels between protein structure prediction and virtual cell efforts. Many of our current approaches to collect data to build a virtual cell model lack the abstraction that links measurement and function. (Link in reply)
2
8
1,617
btw, if you're working on molecular dynamics, drug design, or a related area and are curious about RL and high-performance simulation, please reach out! I'm looking for challenging problems in this space and would love to connect with people who bring domain expertise.
5
8
104
10,234
Replying to @bern_jaeger
Good news for KE:SAI!
3
169