We put a lot of effort into making the experimental setup clean and relatively strict, and hope the sandbox can be useful for future agent evaluation. Also don’t miss the qualitative examples, where the agents came up with some pretty interesting strategies 👀
Hungry for more Astra robot videos? 🤖
How about a large-scale sandboxed evaluation to go with it? We ran 98,000 evaluations across 28 simulated environments and found some surprisingly clever physical reasoning along the way!
New preprint 🧵👇 (1/9)