Excited to release Do As I Do: a pipeline that turns everyday RGB human videos into dexterous robot manipulation trajectories!
Most prior work has been narrow, consisting of just lab recorded demos, egocentric-only, or assuming a closed set of objects. We develop a modular pipeline that can handle Internet, egocentric, exocentric, AND generated videos with virtually any rigid object. Also check out Mahi's post below!
Robots are the bottleneck in scaling robotics, and learning from human video promises to solve it. But how can chaotic human data ever measure up to sanitized, lab-made teleoperation data?
Introducing Do as I Do: establishing a much needed correspondence between human videos and dexterous robot data. Some fun insights below: 🧵