You fine-tune a robot foundation model on a hard task. It gets 25%. Now what?
Introducing Q-Planning, a learning-based harness that lets large black-box robot policies recursively self-improve.
On a hard fine-grained task: 25% → 80% in 100 robot attempts (~30 mins), no extra human data.
q-planning.github.io/
1/n 🧵
Aug 28, 2026 · 11:23 AM UTC
2
4
17
5,212

