Very interesting result: frontier model robot performance did not change with thinking effort.
5 effort levels, 2 simulators, 135 episodes, one variable. 34× the reasoning tokens bought exactly nothing.
Astra is an Incredible step up from GPT 5.6 at robotics tasks
- Astra scored 10/10 vs GPT 5.6 Sol at 3/10
- The step budget is a lot higher since frontier llms are slower.
Astra is very strong at fine-grained spatial reasoning and action sequencing.