To obtain robust autonomy, it has always been the desire for planners to experience large-scale, high-value scenarios in ultrafast closed-loop training with minimal unit cost of data and compute, ideally self-evolving even without offline human demonstration.
TerraZero is our straightforward response at
@AppliedInt to such a desire by scaling self-play reinforcement learning with procedural simulation and zero demonstration:
- Insanely fast in closed loop, i.e., 2.8 million simulation steps per second on 8 GPUs;
- Endlessly streaming hard and diversified corner cases with a procedural generation;
- Robust self-play recipe trading sample efficiency for compute efficiency;
- Zero demonstration in training, and zero-shot emergence across domains (cities, datasets, etc.);
- Minimal unit cost for experiencing high-value data points with a highly scalable framework.
So excited to see that one stack powered by TerraZero can serve various outcomes:
- Setting the state-of-the-art in the challenging, long-tail planning benchmark (InterPlan) with a clear edge;
- Topping the realism among demonstration-free methods in sim agent benchmark (WOSAC);
- Comprehensive practicality with policy covering heterogeneous agents (VRUs, trucks, trailers, etc.) and tackling noisy policy inputs.
On our website, you can also select arbitrary combinations of scenario/agent/view features and see how TerraZero works - play with it!
Really impressed by the potential of large-scale, high-value synthetic data in closed loop powered by ultrafast training, and TerraZero is just the beginning of our Terra Series research - committed to pushing the boundaries of physical AI learned in closed loop with scaled simulation and minimal demonstration. More follow-up work will be announced soon, stay tuned!
TerraZero report:
cas-bridge.xethub.hf.co/xet-…
TerraZero website:
terra-applied.github.io/Terr…
Terra Series website:
terra-applied.github.io/