Introducing the Open Axis Benchmark, a living benchmark engine for honest evaluation of robot manipulation models, built together with
@openroboto.
Robotic models evolve faster every month, while most benchmarks stay frozen. Models overfit to fixed task sets, scores stop reflecting real generalization, and that distorted signal misleads and holds back model evolution.
Open Axis Benchmark isn't just larger. It lets evaluation itself keep pace with the models: each round locks a fresh task set drawn from the growing Axis Library, so models prove they can generalize instead of memorizing the test.
Axis provides the data engine behind it, continuously producing tasks at the scale and diversity needed to genuinely test generalization. With Open Axis Benchmark, the same compounding loop that trains robots also powers how they are evaluated.
Submit your model:
openroboto.ai
Docs:
openroboto.ai/#/docs
See how it works ⬇️